Understand What Broke in Your Voice Agent. Turn by Turn.

Move from a raw transcript to a complete technical breakdown of every call. Trace tool calls, measure millisecond-level latency across STT, LLM, and TTS — and get automated eval scores the moment a call ends.

Get Started

EVENT TIMELINE

A Chronological Map of Every Call Event.

Isolate the exact millisecond where a conversation broke — or where your agent quietly made the wrong call.

Segmented Timeline

Every user utterance, agent response, and tool trigger — laid out in sequence so you can follow exactly what happened and when.

Tool Call Traces

Full raw payloads for every tool call and node transition. No black boxes — see exactly what your agent sent, received, and what came back.

Latency Breakdown

Per-turn performance split by STT, LLM, and TTS provider — so when a call lags, you know exactly which layer is responsible.

Interaction Markers

Interruptions, dead air, and speaker overlaps flagged automatically. See the moments that derailed the conversation before the user hung up.

PER-CALL EVALUATIONS

Run Evals on Every Call. Automatically.

Tuner runs a full evaluation suite the second a call ends. Choose from 30+ pre-built evals covering the most common voice agent scenarios — or write your own checks for what matters to your specific use case.

30+ Pre-Built Evals

Cover the most common voice agent failure modes out of the box — intent detection, compliance rules, hallucinations, response quality, and more. No setup required.

Custom Evals & Behavior Checks

Define your own pass/fail logic for the scenarios that matter to your business. Track what’s unique to your agent, your industry, your users.

Guardrail Alerts

Safety rules applied automatically on every call. The moment your agent crosses a line you’ve defined, you know about it instantly.

Latency & Quality Scores

P50 and P90 latency breakdowns across STT, LLM, and TTS. Longest agent monologue. Talk ratio and pacing. Every number computed automatically — no manual review, no sampling.

THE PHYSICS OF THE CONVERSATION

Per-call voice metrics

Beyond what was said — measure the latency, pacing, and tone that determine whether your agent sounds like a human or a machine.

LATENCY

Time to first byte, split by stage — so you see exactly what slowed the reply.

STT

Speech-to-text: how long until the user’s words are available to the model. Your agent can’t respond until this completes.

LLM / TTFT

Model time to first token after context is ready — your agent’s thinking speed, visible at a glance.

TTS

Text-to-speech: audio generation lag — the gap between a decision and a word the user hears.

TTFB combines the full stack above. Roll up p50 and p90 across all calls — badges mark Slow vs. Good at a glance.

CONVERSATION

Turn-taking, balance, and tone — from the waveform, not from manual scoring.

Talk time

How long the agent held the floor — and whether it left room for the user to speak.

User talk time

Whether callers had space to speak — or gave up trying.

Silence

Dead air that signals confusion, slow responses, or an agent that’s lost.

Crosstalk

Overlapping speech that often signals interruption issues or misread turn-taking.

Longest monologue

Where the agent talked too long — and likely lost the user’s attention.

Sentiment

Overall emotional tone of the conversation, computed automatically — no rubric needed.

Computed on every call automatically — no rubric configuration.