How to handle background noise in a voice agent

Mohamed Salem

Product @ Tuner

Speech-to-text is your voice agent’s ears. It turns what the speaker says into text the LLM can act on. When the speaker is somewhere noisy, those ears start to fail: noise hides parts of words, and nearby voices slip into the transcript. The agent may then look up the wrong booking or act on something the speaker never said.

Noise can also slow replies down. Before answering, the agent has to decide the speaker has finished their turn. Background sound can make it seem like they’re still talking. The speaker stops, the TV behind them keeps going, and the agent keeps waiting.

Background noise is two problems

“Background noise” is really two problems: sounds that cover the speaker’s words, and other people’s voices that get transcribed as if the speaker said them. Most noisy calls have both. A café does, and so does a living room with the TV on. Both can still trip up modern speech-to-text (STT). The second is harder, because to STT, a nearby voice is just more speech.

Environmental noise

A fan humming, keyboard clicks, traffic through an open window. These sounds cover parts of words, so STT can miss or swap a digit in a booking reference. Some STT providers say their models are trained to handle noise, and Deepgram recommends sending audio unaltered. But noise can still cause mistakes, so don’t assume your STT has it covered.

The usual fix is noise suppression (often called noise cancellation), a filter that removes non-speech sound. It helps turn-taking, because a fan or door slam no longer counts as speech. For STT, it can sometimes lower accuracy, so AssemblyAI suggests sending the filtered audio only to voice activity detection (VAD) and turn-taking logic, and the raw audio to STT, where your stack allows it.

Other people’s voices

Now imagine the speaker is discussing that booking while someone nearby says, “Cancel it,” in another conversation. STT may transcribe those words perfectly, and the agent follows the wrong person’s instruction.

Neither fix from the last section helps here. Noise suppression is built to keep speech, and a nearby conversation is speech. And even modern STT can’t tell your speaker’s voice from someone else’s: in Krisp’s tests of 10 STT systems, error rates averaged about 2% on clean audio but 36% on calls with a competing speaker.

The fix: voice isolation

Voice isolation keeps the main speaker’s voice and removes everything else: other people’s voices and environmental noise. It tackles both problems at once, which makes it a better choice than noise suppression for single-speaker lines like booking, support, and sales.

The research backs it up for competing voices. In Krisp’s 2026 benchmark of 265 real recordings across nine STT engines, voice isolation lowered the overall word error rate from 23% to 6%, a 73% drop. On clean phone calls with no second voice, it made results slightly worse, as Krisp itself notes. ai-coustics reports large drops for its own model too. Both are vendors, but a third-party benchmark by SLNG using Deepgram Nova-3 also found voice isolation cut English errors on competing speech by up to 30%.

It runs on the incoming audio before speech-to-text, so STT hears the speaker instead of the room:

Speaker audio → voice isolation → STT → LLM → TTS

It helps turn-taking too. With background voices removed, the agent doesn’t keep waiting through the TV, and nearby chatter is less likely to interrupt it. In Krisp’s tests, voice isolation cut false VAD triggers by 3.5x.

Voice isolation works best when the speaker is the closest, loudest voice on the line. It follows whoever is dominant, not a specific person, so a louder voice nearby can win. Test speakerphone and car calls to see how it copes.

Match it to your channel, too. Phone calls carry less audio detail than web and app calls, so pick an option built for yours. The table below shows which fits where.

If your agent needs to hear more than one person, skip voice isolation: it keeps one voice and treats the rest as background. Use a good STT on raw audio instead, and turn on speaker diarization if the agent needs to know who said what.

Voice isolation providers to consider

We’ve shortlisted voice isolation options worth evaluating. Each one isolates a single primary speaker. Details come from each provider’s documentation (linked in the table), with prices as of September 2026. Pick two or three and compare them on the same calls.

Option

Works with

Built for

Good to know

Krisp VIVA on LiveKit Cloud

LiveKit agents

Web and app calls with krisp.voice_isolation(), SIP calls with krisp.voice_isolation_telephony()

The free plan includes 100 minutes; paid plans include more, then $0.0012/min.

Krisp VIVA Tel

Pipecat (KrispVivaFilter) or the Krisp SDK

Phone calls, up to 16 kHz

Built into Pipecat Cloud: first 10,000 min a month free, then $0.0015/min. Self-hosting needs a Krisp developer account and API key.

Krisp VIVA Pro

Pipecat (KrispVivaFilter) or the Krisp SDK

Web and app calls, up to 32 kHz

The speaker must stay close to the mic. Same Pipecat Cloud pricing as Tel.

ai-coustics Quail Voice Focus

LiveKit Cloud, Pipecat (AICFilter), or the ai-coustics SDK

16 kHz audio, built for voice agents and STT

Comes in small and large versions, so compare both on your hardware. Same price as Krisp VIVA on LiveKit Cloud.

NVIDIA Maxine Speaker Focus

Self-hosted stacks on NVIDIA GPUs

Mono audio, 16 or 48 kHz

Early Access only.

On a managed platform? Turn on its built-in option instead of adding your own filter:

  • Retell: set the denoising mode to “Remove noise + background speech” (adds $0.005/min). Retell notes it can reduce STT accuracy in some cases and struggles on speakerphone or when the speaker is quiet or distant.

  • Vapi: enable Smart Denoising, which is powered by Krisp, with backgroundSpeechDenoisingPlan.smartDenoisingPlan.enabled.

Measure the fix with Tuner

A filter isn’t guaranteed to help, and that includes voice isolation. Retell warns that its own background-speech filter can lower accuracy on calls that are already clean, and even drop short replies like “yes” or “sure.” A filter that fixes noisy calls can hurt quiet ones, so measure the result on your own calls.

1. Before launch: test noisy calls

You could call your voice agent from a café, wait for the coffee grinder, then repeat for every provider. That’s a lot of coffee for a test plan.

A call simulation tool like Tuner makes this repeatable. You can test two ways:

  • Live simulations: Tuner’s AI callers phone your agent and act out realistic conversations. You pick the background (café, street, office, or car) at low, medium, or high volume. The café includes people talking nearby, so one test covers both problems: the noise and the other voices. The background keeps playing for the whole call, even when no one is talking.

  • Replays: Tuner can replay the same recorded caller audio against your agent, so every setup hears exactly the same call. Pick a real call with a background voice, and any change in results comes from your setup, not the conversation.

See our background noise walkthrough for a step-by-step example.

Run the same test four ways, with several calls each:

  1. Quiet, no filter: your baseline.

  2. Noisy, no filter: how much the noise hurts.

  3. Noisy, with voice isolation: whether the fix works. Repeat this run for each provider on your shortlist.

  4. Quiet, with voice isolation: whether the filter hurts calls that were already clean.

Include the moments that break agents most: someone nearby talking while your speaker is silent, and short replies like “yes.” Keep everything else the same across runs: the simulated callers, noise settings, STT model, and prompt. Then check three things:

  • Words: listen to the recording and compare it with the transcript, the text the agent actually received. Check names and numbers first, then look for missing words and words from the background voice.

  • Timing: see how long the agent took to reply and where the time went, such as transcribing or waiting for the speaker to finish. Watch for false interruptions and replies that cut the speaker off.

  • Results: check that the agent did the right thing. Each simulated caller has known details, like a name and a date, so you can check the agent got them right. Tuner’s automatic checks flag wrong details, wrong confirmations, and made-up answers.

In Tuner, each turn shows what the agent heard, how long each step took, and which tools it called (left). Automatic checks then score the call, like flagging a made-up answer (right).

2. After launch: monitor and retest

Keep checking your calls once you’re live. A monitoring tool like Tuner runs the same checks on real calls and sends an alert by email or webhook when failures pile up, such as five slow replies in 15 minutes. A call marked successful can still get a detail wrong, so review flagged calls and a few successful ones too.

Any change can undo your noise fix, so rerun your tests before you release a new filter, STT model, prompt, or turn-detection setting. Replays make this a fair comparison, because the new setup hears exactly the same calls as the old one. When production turns up a new problem, add it to your tests.

💡 Pro tip: Also watch CPU/GPU and memory use alongside call latency. Filtering adds processing work, and an overloaded server can fall behind incoming audio when many calls run at once.

The bottom line

Background noise breaks a voice agent in two ways: it hides the speaker’s words and adds other people’s. Voice isolation tackles both at once, as long as the speaker is the dominant voice. But a filter can also cost accuracy, so don’t ship it on trust. Test it on noisy calls before launch, keep monitoring in production, and rerun the same tests after every change.

Tuner works with Vapi, Retell, LiveKit, Pipecat, Dograh, and custom stacks. Try Tuner and put your agent through café, street, office, and car noise before your customers do.

Speech-to-text is your voice agent’s ears. It turns what the speaker says into text the LLM can act on. When the speaker is somewhere noisy, those ears start to fail: noise hides parts of words, and nearby voices slip into the transcript. The agent may then look up the wrong booking or act on something the speaker never said.

Noise can also slow replies down. Before answering, the agent has to decide the speaker has finished their turn. Background sound can make it seem like they’re still talking. The speaker stops, the TV behind them keeps going, and the agent keeps waiting.

Background noise is two problems

“Background noise” is really two problems: sounds that cover the speaker’s words, and other people’s voices that get transcribed as if the speaker said them. Most noisy calls have both. A café does, and so does a living room with the TV on. Both can still trip up modern speech-to-text (STT). The second is harder, because to STT, a nearby voice is just more speech.

Environmental noise

A fan humming, keyboard clicks, traffic through an open window. These sounds cover parts of words, so STT can miss or swap a digit in a booking reference. Some STT providers say their models are trained to handle noise, and Deepgram recommends sending audio unaltered. But noise can still cause mistakes, so don’t assume your STT has it covered.

The usual fix is noise suppression (often called noise cancellation), a filter that removes non-speech sound. It helps turn-taking, because a fan or door slam no longer counts as speech. For STT, it can sometimes lower accuracy, so AssemblyAI suggests sending the filtered audio only to voice activity detection (VAD) and turn-taking logic, and the raw audio to STT, where your stack allows it.

Other people’s voices

Now imagine the speaker is discussing that booking while someone nearby says, “Cancel it,” in another conversation. STT may transcribe those words perfectly, and the agent follows the wrong person’s instruction.

Neither fix from the last section helps here. Noise suppression is built to keep speech, and a nearby conversation is speech. And even modern STT can’t tell your speaker’s voice from someone else’s: in Krisp’s tests of 10 STT systems, error rates averaged about 2% on clean audio but 36% on calls with a competing speaker.

The fix: voice isolation

Voice isolation keeps the main speaker’s voice and removes everything else: other people’s voices and environmental noise. It tackles both problems at once, which makes it a better choice than noise suppression for single-speaker lines like booking, support, and sales.

The research backs it up for competing voices. In Krisp’s 2026 benchmark of 265 real recordings across nine STT engines, voice isolation lowered the overall word error rate from 23% to 6%, a 73% drop. On clean phone calls with no second voice, it made results slightly worse, as Krisp itself notes. ai-coustics reports large drops for its own model too. Both are vendors, but a third-party benchmark by SLNG using Deepgram Nova-3 also found voice isolation cut English errors on competing speech by up to 30%.

It runs on the incoming audio before speech-to-text, so STT hears the speaker instead of the room:

Speaker audio → voice isolation → STT → LLM → TTS

It helps turn-taking too. With background voices removed, the agent doesn’t keep waiting through the TV, and nearby chatter is less likely to interrupt it. In Krisp’s tests, voice isolation cut false VAD triggers by 3.5x.

Voice isolation works best when the speaker is the closest, loudest voice on the line. It follows whoever is dominant, not a specific person, so a louder voice nearby can win. Test speakerphone and car calls to see how it copes.

Match it to your channel, too. Phone calls carry less audio detail than web and app calls, so pick an option built for yours. The table below shows which fits where.

If your agent needs to hear more than one person, skip voice isolation: it keeps one voice and treats the rest as background. Use a good STT on raw audio instead, and turn on speaker diarization if the agent needs to know who said what.

Voice isolation providers to consider

We’ve shortlisted voice isolation options worth evaluating. Each one isolates a single primary speaker. Details come from each provider’s documentation (linked in the table), with prices as of September 2026. Pick two or three and compare them on the same calls.

Option

Works with

Built for

Good to know

Krisp VIVA on LiveKit Cloud

LiveKit agents

Web and app calls with krisp.voice_isolation(), SIP calls with krisp.voice_isolation_telephony()

The free plan includes 100 minutes; paid plans include more, then $0.0012/min.

Krisp VIVA Tel

Pipecat (KrispVivaFilter) or the Krisp SDK

Phone calls, up to 16 kHz

Built into Pipecat Cloud: first 10,000 min a month free, then $0.0015/min. Self-hosting needs a Krisp developer account and API key.

Krisp VIVA Pro

Pipecat (KrispVivaFilter) or the Krisp SDK

Web and app calls, up to 32 kHz

The speaker must stay close to the mic. Same Pipecat Cloud pricing as Tel.

ai-coustics Quail Voice Focus

LiveKit Cloud, Pipecat (AICFilter), or the ai-coustics SDK

16 kHz audio, built for voice agents and STT

Comes in small and large versions, so compare both on your hardware. Same price as Krisp VIVA on LiveKit Cloud.

NVIDIA Maxine Speaker Focus

Self-hosted stacks on NVIDIA GPUs

Mono audio, 16 or 48 kHz

Early Access only.

On a managed platform? Turn on its built-in option instead of adding your own filter:

  • Retell: set the denoising mode to “Remove noise + background speech” (adds $0.005/min). Retell notes it can reduce STT accuracy in some cases and struggles on speakerphone or when the speaker is quiet or distant.

  • Vapi: enable Smart Denoising, which is powered by Krisp, with backgroundSpeechDenoisingPlan.smartDenoisingPlan.enabled.

Measure the fix with Tuner

A filter isn’t guaranteed to help, and that includes voice isolation. Retell warns that its own background-speech filter can lower accuracy on calls that are already clean, and even drop short replies like “yes” or “sure.” A filter that fixes noisy calls can hurt quiet ones, so measure the result on your own calls.

1. Before launch: test noisy calls

You could call your voice agent from a café, wait for the coffee grinder, then repeat for every provider. That’s a lot of coffee for a test plan.

A call simulation tool like Tuner makes this repeatable. You can test two ways:

  • Live simulations: Tuner’s AI callers phone your agent and act out realistic conversations. You pick the background (café, street, office, or car) at low, medium, or high volume. The café includes people talking nearby, so one test covers both problems: the noise and the other voices. The background keeps playing for the whole call, even when no one is talking.

  • Replays: Tuner can replay the same recorded caller audio against your agent, so every setup hears exactly the same call. Pick a real call with a background voice, and any change in results comes from your setup, not the conversation.

See our background noise walkthrough for a step-by-step example.

Run the same test four ways, with several calls each:

  1. Quiet, no filter: your baseline.

  2. Noisy, no filter: how much the noise hurts.

  3. Noisy, with voice isolation: whether the fix works. Repeat this run for each provider on your shortlist.

  4. Quiet, with voice isolation: whether the filter hurts calls that were already clean.

Include the moments that break agents most: someone nearby talking while your speaker is silent, and short replies like “yes.” Keep everything else the same across runs: the simulated callers, noise settings, STT model, and prompt. Then check three things:

  • Words: listen to the recording and compare it with the transcript, the text the agent actually received. Check names and numbers first, then look for missing words and words from the background voice.

  • Timing: see how long the agent took to reply and where the time went, such as transcribing or waiting for the speaker to finish. Watch for false interruptions and replies that cut the speaker off.

  • Results: check that the agent did the right thing. Each simulated caller has known details, like a name and a date, so you can check the agent got them right. Tuner’s automatic checks flag wrong details, wrong confirmations, and made-up answers.

In Tuner, each turn shows what the agent heard, how long each step took, and which tools it called (left). Automatic checks then score the call, like flagging a made-up answer (right).

2. After launch: monitor and retest

Keep checking your calls once you’re live. A monitoring tool like Tuner runs the same checks on real calls and sends an alert by email or webhook when failures pile up, such as five slow replies in 15 minutes. A call marked successful can still get a detail wrong, so review flagged calls and a few successful ones too.

Any change can undo your noise fix, so rerun your tests before you release a new filter, STT model, prompt, or turn-detection setting. Replays make this a fair comparison, because the new setup hears exactly the same calls as the old one. When production turns up a new problem, add it to your tests.

💡 Pro tip: Also watch CPU/GPU and memory use alongside call latency. Filtering adds processing work, and an overloaded server can fall behind incoming audio when many calls run at once.

The bottom line

Background noise breaks a voice agent in two ways: it hides the speaker’s words and adds other people’s. Voice isolation tackles both at once, as long as the speaker is the dominant voice. But a filter can also cost accuracy, so don’t ship it on trust. Test it on noisy calls before launch, keep monitoring in production, and rerun the same tests after every change.

Tuner works with Vapi, Retell, LiveKit, Pipecat, Dograh, and custom stacks. Try Tuner and put your agent through café, street, office, and car noise before your customers do.

Tuner helps voice AI teams test, monitor, and debug calls before issues reach users.