Understand What Broke in Your Voice Agent. Turn by Turn.
Move from a raw transcript to a complete technical breakdown of every call. Trace tool calls, measure millisecond-level latency across STT, LLM, and TTS — and get automated eval scores the moment a call ends.
Get Started
EVENT TIMELINE
A Chronological Map of Every Call Event.
Isolate the exact millisecond where a conversation broke — or where your agent quietly made the wrong call.
✓
Segmented Timeline
Every user utterance, agent response, and tool trigger — laid out in sequence so you can follow exactly what happened and when.
✓
Tool Call Traces
Full raw payloads for every tool call and node transition. No black boxes — see exactly what your agent sent, received, and what came back.
✓
Latency Breakdown
Per-turn performance split by STT, LLM, and TTS provider — so when a call lags, you know exactly which layer is responsible.
✓
Interaction Markers
Interruptions, dead air, and speaker overlaps flagged automatically. See the moments that derailed the conversation before the user hung up.

PER-CALL EVALUATIONS
Run Evals on Every Call. Automatically.
Tuner runs a full evaluation suite the second a call ends. Choose from 30+ pre-built evals covering the most common voice agent scenarios — or write your own checks for what matters to your specific use case.

✓
30+ Pre-Built Evals
Cover the most common voice agent failure modes out of the box — intent detection, compliance rules, hallucinations, response quality, and more. No setup required.
✓
Custom Evals & Behavior Checks
Define your own pass/fail logic for the scenarios that matter to your business. Track what’s unique to your agent, your industry, your users.
✓
Guardrail Alerts
Safety rules applied automatically on every call. The moment your agent crosses a line you’ve defined, you know about it instantly.
✓
Latency & Quality Scores
P50 and P90 latency breakdowns across STT, LLM, and TTS. Longest agent monologue. Talk ratio and pacing. Every number computed automatically — no manual review, no sampling.
THE PHYSICS OF THE CONVERSATION
Per-call voice metrics
Beyond what was said — measure the latency, pacing, and tone that determine whether your agent sounds like a human or a machine.
LATENCY
Time to first byte, split by stage — so you see exactly what slowed the reply.
STT
Speech-to-text: how long until the user’s words are available to the model. Your agent can’t respond until this completes.
LLM / TTFT
Model time to first token after context is ready — your agent’s thinking speed, visible at a glance.
TTS
Text-to-speech: audio generation lag — the gap between a decision and a word the user hears.
TTFB combines the full stack above. Roll up p50 and p90 across all calls — badges mark Slow vs. Good at a glance.
CONVERSATION
Turn-taking, balance, and tone — from the waveform, not from manual scoring.
Talk time
How long the agent held the floor — and whether it left room for the user to speak.
User talk time
Whether callers had space to speak — or gave up trying.
Silence
Dead air that signals confusion, slow responses, or an agent that’s lost.
Crosstalk
Overlapping speech that often signals interruption issues or misread turn-taking.
Longest monologue
Where the agent talked too long — and likely lost the user’s attention.
Sentiment
Overall emotional tone of the conversation, computed automatically — no rubric needed.
Computed on every call automatically — no rubric configuration.
