Full observability for your Pipecat voice agents

Add the Tuner observer to your Pipecat pipeline and capture every call's transcript, latency, usage, and cost automatically, so you catch hallucinations, broken flows, and missed intents before your callers do.

from tuner_pipecat_sdk import Observer

observer = Observer(

    api_key=TUNER_API_KEY,

    workspace_id=42,

    agent_id="my-agent",

    call_id=str(uuid4()),

)


# drop it into your pipeline, right after TTS

pipeline = Pipeline([..., tts, transport.output()])

task = PipelineTask(

pipeline,

params=PipelineParams(

enable_metrics=True,

enable_usage_metrics=True,

),

observers=[observer, observer.latency_observer, turn_tracker],

)

Setup

Integrate in under two minutes

No re-architecting your pipeline. The Tuner observer attaches to your existing Pipecat agent and starts capturing production data immediately.

01

Install the SDK

pip install tuner-pipecat-sdk. Works with pipecat-ai 1.0+ on Python 3.11–3.13. Add the flows extra if you run pipecat-flows.

02

Set your Credentials

Drop in your Tuner API key, workspace ID, and agent ID — via environment variables or inline in code.

03

Create the observer

Add Observer for a plain pipeline, or FlowsObserver for pipecat-flows. Pass your Tuner API key, workspace ID, and agent ID.

04

Monitor calls in Tuner

Transcripts, latency, usage, and cost flow into your dashboard automatically — no manual API calls, ready to analyze and monitor.

Features

Everything you need to run Pipecat agents in production

Turn production from a black box into something you can actually monitor, measure, and improve.

Catch failures early

Automatically flag hallucinations, broken flows, dead air, early hangups, and other failure conditions before they show up in your churn data.

See where latency comes from

Break out STT, TTS, and LLM latency at p50 and p90, so you can see exactly which part of the voice stack is slowing conversations down.

Get alerted when something breaks

Get notified when red flags, failed evals, or other conditions appear in production — while there’s still time to fix them.

Simulate calls before you ship

Stress-test your agent over SIP before launch and after every change, using the same evals that monitor your live traffic.

LangGraph & LangChain capture

Record LangGraph and LangChain node transitions, tool calls, and timing alongside session data, so you can see what your logic layer was doing during the call.

Track cost on every call

Attach a cost calculator and track LLM, TTS, and STT spend on every session — no separate billing pipeline required.

Why Tuner

See what's happening in production

Voice agents fail quietly, and at a scale no team can review by hand. Tuner turns every production call into structured data you can search, debug, evaluate, alert on, and test against.

01

Debug in minutes

When a call goes wrong, see exactly what happened and where — the full transcript, every turn, stage-level latency, tool calls, and conversation state. No more guessing from sparse logs.

02

Get alerted the moment something breaks

Don't wait for a customer complaint. Create alerts using red flags, metrics, evals, and multiple conditions, then get notified when a critical call fails or a problem starts appearing across your traffic.

03

Test before you ship

Run realistic call simulations over SIP before launch and after every change. Score simulations with the same evals you use in production, so regressions are caught before your callers find them.

04

Diagnose failures at scale

When you're handling thousands of calls a day, manual review doesn't scale. Tuner finds patterns across your traffic, traces failures to their source, and helps you turn what you learn into evals and alerts.

05

Measure what matters

Track outcomes, intents, extracted data, latency, cost, and custom evals across your Retell traffic. Filter down to the calls that matter and understand how your agent is performing over time.

01

Debug in minutes

When a call goes wrong, see exactly what happened and where — the full transcript, every turn, stage-level latency, tool calls, and conversation state. No more guessing from sparse logs.

02

Get alerted when something breaks

Don't wait for a customer complaint. Create alerts using red flags, metrics, evals, and multiple conditions, then get notified when a critical call fails or a problem starts appearing across your traffic.

05

Measure what matters

Track outcomes, intents, extracted data, latency, cost, and custom evals across your Retell traffic. Filter down to the calls that matter and understand how your agent is performing over time.

03

Test before you ship

Run realistic call simulations over SIP before launch and after every change. Score simulations with the same evals you use in production, so regressions are caught before your callers find them.

04

Diagnose failures at scale

When you're handling thousands of calls a day, manual review doesn't scale. Tuner finds patterns across your traffic, traces failures to their source, and helps you turn what you learn into evals and alerts.

Comparison

Tuner vs Pipecat Evals

Pipecat Evals is built for development — fast, local behavioral tests you run in CI. Tuner is the production layer that scores real calls after you ship. Most teams run both.

Capability

Tuner

Pipecat

Vendor-independent observability, eliminating the conflict of a platform evaluating its own output

Evals pricing built for scale: tuner price per call, no per minute surcharge

Built-in flags (hallucination, dead air, early hangup)

Root-cause diagnosis with a specific fix, not just metrics

30+ voice quality metrics & red flags out of the box

Drift & regression alerts over time

SIP call simulations with AI agents, using your live evals

Turn-by-turn transcripts & latency traces

FAQ

Frequently asked questions

Common questions about connecting Tuner to your Pipecat agents.

Read the docs

Which Pipecat versions are supported?

Do I have to restructure my pipeline?

How long does setup take?

What gets captured?

Can I test my agent before going live?

Does it work with SIP / phone calls?

Does Tuner support alerts and monitoring?

Can I define my own evaluations and metrics?

How is Tuner priced?

Is my call data private and secure?