Know when your voice agent is actually working

Monitor every production conversation. Evaluate every behavior that matters. Catch failures before they become patterns.

Tuner · Monitoring1 firing
Last 24h1,284 calls6 rules active
Alerts fired
14+9 from yesterday
Red flag rate
2.8%4x on refunds
Policy grounding
93.2%−1.4 pts
TTFB p90
3.37spast 3.0s budget
Alert feednewest first
14:02
hallucination_watchcritical
Policy grounding failed while the caller was still on the line.call_2a63c9 · sentiment 2.1
11:47
latency_budgetdrift
TTFB p90 drifted from 2.91s to 3.37s over six hours.threshold · 214 calls affected
09:15
refund_regressionpattern
Red flag volume up 4x on refund intent since v2.4.1.36 calls · prompt v2.4.1
08:40
identity_verificationresolved
5 consecutive failures cleared after the tool timeout was fixed.acknowledged by dana@
hallucination_watch · firing now
14:02:18 · 40s ago
WHEN eval policy_grounding = fail
AND sentiment < 3
Cited evidence · call_2a63c9
01:04"I bought it about two months ago and I'd like to send it back."
01:12"You can return it any time within ninety days."
retrieved returns.window = 30 days
Signals on this call
Policy groundingFail
Sentiment2.1
Intent capturedPass
Latency budget4.8s
slack #voice-opswebhook /alertsemail · on-call
Failed policy_grounding · per hourdeploy v2.4.1 at 08:00
02:0008:0014:00
Tuner · Monitoring1 firing
Last 24h1,284 calls6 rules active
Alerts fired
14+9 from yesterday
Red flag rate
2.8%4x on refunds
Policy grounding
93.2%−1.4 pts
TTFB p90
3.37spast 3.0s budget
Alert feednewest first
14:02
hallucination_watchcritical
Policy grounding failed while the caller was still on the line.call_2a63c9 · sentiment 2.1
11:47
latency_budgetdrift
TTFB p90 drifted from 2.91s to 3.37s over six hours.threshold · 214 calls affected
09:15
refund_regressionpattern
Red flag volume up 4x on refund intent since v2.4.1.36 calls · prompt v2.4.1
08:40
identity_verificationresolved
5 consecutive failures cleared after the tool timeout was fixed.acknowledged by dana@
hallucination_watch · firing now
14:02:18 · 40s ago
WHEN eval policy_grounding = fail
AND sentiment < 3
Cited evidence · call_2a63c9
01:04"I bought it about two months ago and I'd like to send it back."
01:12"You can return it any time within ninety days."
retrieved returns.window = 30 days
Signals on this call
Policy groundingFail
Sentiment2.1
Intent capturedPass
Latency budget4.8s
slack #voice-opswebhook /alertsemail · on-call
Failed policy_grounding · per hourdeploy v2.4.1 at 08:00
02:0008:0014:00

Works With

Custom Stack

Works With

Custom Stack

Works With

Custom Stack

Evals & Guardrails

Define what "good" looks like. Score every call against it.

Define your criteria in plain language. Tuner scores every production call automatically — no sampling, no manual review. Run the same evals in simulation and production, so what you test is what you monitor.

30+ pre-built evals.

Cover common voice-agent failure modes out of the box, including hallucinations, scope boundaries, escalation handling, and out-of-hours behavior.

Custom Evals

Define the pass/fail criteria that matter to your business. Tuner scores every call against your rules.

Dynamic evals

Score each call against the instructions and context the agent actually received, with different expectations for each customer, session, or task.

See why it passed or failed

Every eval result includes the transcript evidence behind the verdict, so you can see exactly what happened instead of trusting an unexplained score.

Tuner · Evals30 pre-built · 4 custom
policy_grounding
customproduction + simulation
"The agent must only state return windows that appear in the retrieved policy document. Quoting a window from memory is a fail."
Runs on every callScored 1,284 of 1,284Pass rate 93.2%
Result · call_2a63c9Fail
Cited evidence
01:04"I bought it about two months ago and I'd like to send it back."
01:12"You can return it any time within ninety days."
Retrieved policyreturns.window = 30 days from delivery
Intent capturedPass
TonePass
Latency budget4.8s

Alerts

Know the moment something breaks

Catch problems as they happen instead of waiting for a report. Tuner alerts you on individual failures and patterns across your traffic.

Alert on any signal

Trigger alerts from intents, outcomes, extracted data, red flags, evals, or any metric Tuner computes. If Tuner can detect it, you can alert on it.

User-defined conditions

Build alert rules from multiple signals. For example: alert when hallucination is detected and sentiment drops below 3, or when 5+ calls fail identity verification within 10 minutes.

Catch failures and patterns

Alert immediately on a critical call, or when a problem starts appearing across calls — rising latency, increasing red flags, or falling resolution rates.

Route alerts where your team works

Send alerts to email, webhooks, or Slack.

Tuner · Alerts6 rules active
Rule · hallucination_watchenabled
WHEN eval policy_grounding = fail
AND sentiment < 3
OR 5+ calls in 10 min fail identity_verification
Route toslack #voice-opswebhook /alertsemail · on-call
Fired today3 alerts
14:02
Policy grounding failed on a live callcall_2a63c9 · per-call · sentiment 2.1
critical
11:47
TTFB p90 drifted past 3.0sthreshold · 2.91s → 3.37s over 6h
drift
09:15
Red flag volume up 4x on refund intentpattern · 36 calls · since v2.4.1
pattern

Call Analysis

Understand every conversation automatically.

Extract the signals you need from every conversation — without manual labeling or sampling.

Know how every call ended

Detect outcomes such as resolution, escalation, abandonment, or incorrect action automatically.

Know what callers are asking for

Classify caller intent automatically and group conversations by type — Sales, Support, Onboarding, Billing, and more.

Extract the data you need

Pull dates, numbers, names, and custom entities from conversations and store them as queryable metadata.

Tuner · Call analysis1,284 calls · last 30 days
Outcomesauto-labeled
Resolved
71.4%
Escalated
13.8%
Abandoned
8.9%
Incorrect action
5.9%
Intents detected
Support
412
Billing
288
Sales
201
Onboarding
96
Extracted fieldscall_2a63c9
order_idORD-48213
purchased_on2026-06-14
refund_amount$128.40
customer_tierplus

Red Flags

Auto-tag failures the moment they happen.

Some failures can't be captured by a generic metric. Define your own rules in plain language and automatically tag the calls that match.

Build rules from any signal

Create rules using evals, metrics, extracted data, call metadata, or combinations of signals.

Surface high-risk calls

Flag calls where multiple signals point to a serious issue, so your team can focus on the conversations that need attention.

Create contextual labels

Tag patterns you want to track without treating them as failures.

Tuner · Red flags36 flagged · last 7 days
Rulematched 36 of 1,284
"Flag any call where the agent quotes a policy number that isn't in the retrieved document."
eval: policy_groundingmetric: sentimentmetadata: intent
High-risk queueby signal count
call_2a63c9
hallucinationpolicy failsentiment drop
3
call_2a63d1
tool timeoutlatency 4.8s
2
call_2a63e4
identity unverified
1
Contextual labels · tracked, not failures
competitor mentioned · 41discount requested · 27needs review · 12

Connect

Get connected in five minutes or less.

No-code

~2 mins

Retell, Vapi, Dograh. Add a key. Calls sync automatically.

SDKs

~5-10 mins

LiveKit, Pipecat, or your own Python stack. Two lines, and you get timing your platform can't give you.

API

~30 mins

Any language, any voice stack.

Connect

Get connected in five minutes or less.

No-code

~2 mins

Retell, Vapi, Dograh. Add a key. Calls sync automatically.

SDKs

~5-10 mins

LiveKit, Pipecat, or your own Python stack. Two lines, and you get timing your platform can't give you.

API

~30 mins

Any language, any voice stack.

Connect

Get connected in five minutes or less.

No-code

~2 mins

Retell, Vapi, Dograh. Add a key. Calls sync automatically.

SDKs

~5-10 mins

LiveKit, Pipecat, or your own Python stack. Two lines, and you get timing your platform can't give you.

API

~30 mins

Any language, any voice stack.