← ALL EVENTS

Builders of Voice AI · Episode 2

On-Prem Voice Agents: Why Enterprises Are Choosing to Own Their Stack

On-Prem Voice Agents: Why Enterprises Are Choosing to Own Their Stack

On-Prem Voice Agents: Why Enterprises Are Choosing to Own Their Stack

Mai Medhat sits down with Dograh co-founders Pritesh Kumar (CEO) and Abhishek Kumar (CTO) to unpack the real cost, compliance, and reliability math behind self-hosting voice AI.

Mai Medhat sits down with Dograh co-founders Pritesh Kumar (CEO) and Abhishek Kumar (CTO) to unpack the real cost, compliance, and reliability math behind self-hosting voice AI.

July 21, 2026

/

~1 hour

/

Mai Medhat, Pritesh Kumar & Abhishek Kumar

About this episode

Dograh is the open-source, self-hostable voice platform, and since launch it's crossed 5,000 GitHub stars and 1,100+ forks. Pritesh and Abhishek join Mai to get into why enterprises are choosing to own their voice AI stack instead of renting it: the compliance stories driving the decision, the real cost math behind self-hosting (and where it breaks even), and the reliability problem every voice AI builder eventually runs into.

10 things we learned

01 · Data sovereignty is still the #1 reason enterprises self-host A global enterprise in Taiwan self-hosts everything on Dograh because they won't let call data leave the country. An Australian BPO wanted full air-gapped, on-prem GPUs. They wouldn't even share a call script with ChatGPT. Compliance isn't a checkbox for these teams, it's the whole decision.

02 · Self-hosting isn't cheaper, it's differently priced Running a voice agent fully on-prem at ~5 concurrent calls costs roughly $500/month: about $300 for a GPU running a small model, $150–300 for the orchestration server. The real cost isn't the infra, it's the DevOps effort to keep it running, something almost everyone underestimates.

03 · There's a real volume threshold for self-hosting Below it, don't bother. Dograh's rule of thumb: self-hosting starts to make sense around 100K–200K calls a month. Below that, or if speed to market matters more than control, a cloud platform wins every time.

04 · Sometimes it's not a logical decision, it's an emotional one "It's like buying a house instead of renting. It doesn't always make financial sense, but people do it anyway." Pritesh described a Florida restaurant chain whose founder self-hosted against their advice, purely because he loves building. Passion for owning your stack is a real, recurring driver.

05 · The reliability math nobody talks about Multiply it out: ~80% transcription accuracy × ~90% LLM accuracy × ~90% speech generation accuracy ≈ 63% odds of a clean turn. Over a 10-turn conversation, that's single-digit odds of a flawless call. It's the clearest explanation yet for why voice AI feels so much harder to get right than it looks.

06 · The middle ground: data-boundary hosting, not full air-gap Most enterprises don't need to fully air-gap. Hosting models inside a cloud provider's VPC or data perimeter keeps data from ever leaving the boundary, satisfying compliance without the cost and complexity of running GPUs in a basement.

07 · Post-launch is where the real surprises show up Mai shared a real estate voice agent built to answer listing questions, but once live, users mostly asked whether a neighborhood was safe or good for schools. The agent had no context for that, so it hallucinated confident-sounding answers instead. The gap between what you build for and what people actually ask is where most failures start.

08 · Being open source makes them faster, not just cheaper Dograh's users push PRs and flag gaps directly in their community: feedback closed platforms simply don't get. Pritesh's blunt (and funny) answer for why they went open source in the first place: "we wanted to watch the world burn." The real reason is more standard OSS conviction, but the line got a laugh for good reason.

09 · Even a fully open-source platform didn't build its own observability Dograh ships as a complete stack (database, file storage, cache, its own MCP), but for monitoring, they integrate Tuner rather than building it themselves. Abhishek: you don't fiddle with SDKs, "you literally go and drop a box, create an account on Tuner, configure the credential, and you're set."

10 · Speech-to-speech vs. cascade isn't settled, it's use-case specific Speech-to-speech wins on multilingual conversations and latency (Gemini's live models handle 70+ languages interchangeably). Cascade wins when you need tight control, like outbound calls with strict logic. Pritesh's take: build for speech-to-speech if you're building for where the industry is headed.

Standout moments

"You're looking at .63 to the power of 10: that's the probability of a call going right." The single most quotable moment of the episode. Pritesh's back-of-envelope math on why voice agents fail more than people expect: compounding errors across transcription, reasoning, and speech generation, turn after turn.

The Australian BPO that wouldn't trust ChatGPT with a call script Full data sovereignty, on-prem GPUs, and a decision-maker who wouldn't even paste an internal script into an AI model to ask for suggestions. A vivid, real example of how far compliance sensitivity goes in regulated industries.

"Self-hosting isn't just a commercial decision, it's an emotional one." The house-buying analogy: renting is often the logical choice, but people buy anyway. Explains a pattern Dograh sees constantly: teams self-hosting well before it makes financial sense, purely out of conviction.

"We wanted to watch the world burn." Pritesh's tongue-in-cheek answer for why they built Dograh open source, delivered with a laugh, but underneath it, a real point about how open source lets closed-source incumbents get out-iterated.

The Tuner integration, in the guests' own words Not a pitch: a credible team explaining a build-vs-buy call. Abhishek: full observability without touching a codebase. Mai: "this is probably the easiest integration we have."

Guests

Mai Medhat — Founder and CEO, Tuner
Building the reliability layer for voice AI: observability, evals, and automated testing.

Pritesh Kumar — Co-Founder, Dograh
YC alum and exit founder building the open-source, self-hostable alternative to Vapi and Retell.

Abhishek Kumar — Co-Founder, Dograh
15+ years building products; leads the engineering behind Dograh's self-hosted voice orchestration.

Never Miss an Event

Never Miss an Event

Never Miss an Event

Get notified when we announce new episodes and upcoming live conversations with voice AI leaders.