Launching Replays and Datasets for Voice Agents

Mai Medhat

CEO & Co-founder at Tuner

Today we're launching Replays and Datasets in Tuner. They solve two problems every team building voice agents runs into: you can't reproduce the call that failed, and you can't pick your stack from someone else's benchmark.

Replays and Datasets fix both. Save real calls into a Dataset, then replay them against your agent as many times as you need, with the exact same caller audio every time.

The call you can't make twice

Voice agents rarely fail on the calls you tested. They fail on the ones you didn't see coming:

  • A caller on speakerphone in a moving car, and the agent hears "Tuesday" instead of "Thursday."

  • An accent your STT struggles with, so the agent keeps asking the caller to repeat themselves.

  • A caller who reads a date in a way your agent didn't expect, and the booking tool call fails.

  • A caller who talks over the agent, and turn detection cuts them off mid-sentence.

Tuner already flags the Failure and shows you where the call broke. The hard part is proving your fix works on this exact call. Calling the agent yourself from a quiet office doesn't prove it. AI Callers and Call Simulation get close, but they can't reproduce the exact voice, noise, and timing that broke it. The only real test is the call that failed. That's why we built Replays.

How Replays work

  1. Build a Dataset. Pick calls from your Tuner call logs, or upload recordings (wav, mp3, or m4a). We separate the speakers, you pick the caller, and we save their turns.

  2. Change your agent. Fix the prompt, swap the model, update the stack.

  3. Run a Replay. We call your agent and play back the same caller turns, like a live call.

  4. See the result. Every Replay is scored against the Evals, Call Outcomes, and Intents you already use in production, with latency measured on every run.

From Failure to verified fix, on the same call

Here's what this looks like in practice. A dental clinic runs a voice agent that books appointments. A patient calls from their car, on speakerphone, with the road loud in the background, and asks to move their appointment to "Friday the fourteenth." The agent mishears the date, the reschedule tool call fails, and the call ends with no new booking.

1. Spot the Failure. Tuner flags the call and points at the turn where the tool call errored.

2. Save the call. Add it to a Dataset straight from your call logs. Road noise and all.

3. Reproduce it. Replay it against your current agent before changing anything. If the Failure shows up again, it's a real problem, not a one-off glitch, and you have a test that fails for the right reason.

4. Fix the agent. A prompt change, a different STT, normalising the date before the tool call. Whatever you think fixes it.

5. Replay it again. Same patient, same car, same "Friday the fourteenth." This time the date is right, the tool call succeeds, and the call passes all evals.

You're not hoping the fix worked. You watched it work on the call that broke.

Your Datasets become your test suite

That call stays in your Dataset, and so does every hard call you add: the noisy ones, the strong accents, the interruptions, the edge cases. Replay the whole Dataset with every prompt update, model swap, or deployment. If an old Failure comes back, you'll see it before your callers hear it.

Pick your stack on your own calls

There's a new model every week. It tops the leaderboard, LinkedIn goes crazy about it, and a few days later another one takes its place.

A leaderboard can't tell you if that model is right for your agent. Public benchmarks run on generic audio that may overlap with what the models were trained on. Your callers are specific: their language, their accents, their background noise, your industry's words. And a model never runs alone. Turn detection, orchestration, noise handling, and speaker detection can change how the same model performs and the latency your caller actually feels.

So test it on your calls, in your stack. Swap in the new STT or LLM, replay your Dataset, and compare the results side by side with your current setup. Same calls, same scoring, latency measured on every run.

That's your own benchmark. Built on your data, for your use case, and it doesn't change every week.

Get started

Replays and Datasets are live in Tuner today. If you want to try them, start with a call that broke last week: add it to a Dataset from your call logs, or upload the recording, and replay it. You don't need any live traffic to try it.

Try it on your own calls


We built this because we believe building voice agents should feel like building software. When something breaks, you reproduce it. When you fix it, you prove it. And every edge case you've hit lives in a test suite that runs on every change.

Today we're one step closer to that.

Today we're launching Replays and Datasets in Tuner. They solve two problems every team building voice agents runs into: you can't reproduce the call that failed, and you can't pick your stack from someone else's benchmark.

Replays and Datasets fix both. Save real calls into a Dataset, then replay them against your agent as many times as you need, with the exact same caller audio every time.

The call you can't make twice

Voice agents rarely fail on the calls you tested. They fail on the ones you didn't see coming:

  • A caller on speakerphone in a moving car, and the agent hears "Tuesday" instead of "Thursday."

  • An accent your STT struggles with, so the agent keeps asking the caller to repeat themselves.

  • A caller who reads a date in a way your agent didn't expect, and the booking tool call fails.

  • A caller who talks over the agent, and turn detection cuts them off mid-sentence.

Tuner already flags the Failure and shows you where the call broke. The hard part is proving your fix works on this exact call. Calling the agent yourself from a quiet office doesn't prove it. AI Callers and Call Simulation get close, but they can't reproduce the exact voice, noise, and timing that broke it. The only real test is the call that failed. That's why we built Replays.

How Replays work

  1. Build a Dataset. Pick calls from your Tuner call logs, or upload recordings (wav, mp3, or m4a). We separate the speakers, you pick the caller, and we save their turns.

  2. Change your agent. Fix the prompt, swap the model, update the stack.

  3. Run a Replay. We call your agent and play back the same caller turns, like a live call.

  4. See the result. Every Replay is scored against the Evals, Call Outcomes, and Intents you already use in production, with latency measured on every run.

From Failure to verified fix, on the same call

Here's what this looks like in practice. A dental clinic runs a voice agent that books appointments. A patient calls from their car, on speakerphone, with the road loud in the background, and asks to move their appointment to "Friday the fourteenth." The agent mishears the date, the reschedule tool call fails, and the call ends with no new booking.

1. Spot the Failure. Tuner flags the call and points at the turn where the tool call errored.

2. Save the call. Add it to a Dataset straight from your call logs. Road noise and all.

3. Reproduce it. Replay it against your current agent before changing anything. If the Failure shows up again, it's a real problem, not a one-off glitch, and you have a test that fails for the right reason.

4. Fix the agent. A prompt change, a different STT, normalising the date before the tool call. Whatever you think fixes it.

5. Replay it again. Same patient, same car, same "Friday the fourteenth." This time the date is right, the tool call succeeds, and the call passes all evals.

You're not hoping the fix worked. You watched it work on the call that broke.

Your Datasets become your test suite

That call stays in your Dataset, and so does every hard call you add: the noisy ones, the strong accents, the interruptions, the edge cases. Replay the whole Dataset with every prompt update, model swap, or deployment. If an old Failure comes back, you'll see it before your callers hear it.

Pick your stack on your own calls

There's a new model every week. It tops the leaderboard, LinkedIn goes crazy about it, and a few days later another one takes its place.

A leaderboard can't tell you if that model is right for your agent. Public benchmarks run on generic audio that may overlap with what the models were trained on. Your callers are specific: their language, their accents, their background noise, your industry's words. And a model never runs alone. Turn detection, orchestration, noise handling, and speaker detection can change how the same model performs and the latency your caller actually feels.

So test it on your calls, in your stack. Swap in the new STT or LLM, replay your Dataset, and compare the results side by side with your current setup. Same calls, same scoring, latency measured on every run.

That's your own benchmark. Built on your data, for your use case, and it doesn't change every week.

Get started

Replays and Datasets are live in Tuner today. If you want to try them, start with a call that broke last week: add it to a Dataset from your call logs, or upload the recording, and replay it. You don't need any live traffic to try it.

Try it on your own calls


We built this because we believe building voice agents should feel like building software. When something breaks, you reproduce it. When you fix it, you prove it. And every edge case you've hit lives in a test suite that runs on every change.

Today we're one step closer to that.

Tuner helps voice AI teams test, monitor, and debug calls before issues reach users.