# Deepgram Voice Agent API — The speech-infrastructure vendor's own voice-agent API. Managed STT, LLM and TTS over a single WebSocket, $4.50/hr flat, with 6 LLM providers and self-hosted or VPC deployment on the same product.

> Source: The Agents Index — https://theagentsindex.com/deepgram-voice-agent-api (structured, researched, re-verified)
> Facts last verified: 2026-08-23

Deepgram Voice Agent API is the speech-infrastructure company's own product for building real-time voice agents. A single WebSocket connection bundles speech-to-text, LLM reasoning and text-to-speech with built-in barge-in detection, turn-taking prediction and mid-conversation function calling. Deepgram was founded in 2015 by three physicists from the University of Michigan, has sold STT and TTS infrastructure to 1,300+ organizations including NASA, Spotify and Citibank since going through Y Combinator in 2016, and raised a $130 million Series C in January 2026 at a $1.3 billion valuation. The Voice Agent API is Deepgram's own orchestration layer, built on that same underlying speech infrastructure.

| Fact | Value |
| --- | --- |
| Website | https://deepgram.com/product/voice-agent-api |
| Pricing | Flat $4.50/hr on the Standard tier, billed on WebSocket connection time. Pay-as-you-go rates range from $0.050/min (Custom, fully bring-your-own LLM and TTS) up to $0.163/min (Advanced with Deepgram's own STT and TTS), with promotional cuts through September 12, 2026. New accounts get a one-time $200 credit, no credit card required, and Flux TTS is free inside the Voice Agent through the same window. An annual Growth Plan of pre-paid credits offers up to 20% off the standard rates. Enterprise deployments (including dedicated single-tenant or fully self-hosted) are custom-quoted. |
| API | Yes |
| Best for | Engineering teams that want a managed speech pipeline with real LLM-provider choice, and regulated enterprises that need a self-hosted or VPC deployment option a typical closed-source voice-agent vendor does not offer. |
| Not for | Teams that want a configured, no-code voice-agent product with a dashboard out of the box. Synthflow AI, Retell AI or Bland AI fit that need better. Deepgram Voice Agent API is infrastructure, not a turnkey product. |

## Pricing

| Tier | Price |
| --- | --- |
| Pay-As-You-Go | $4.50/hr Standard / Standard to Custom (BYO LLM and TTS) |
| Growth Plan | $0.068 – $0.041 / min / pre-paid annual credits, up to 20% off PAYG |
| Enterprise | Custom (contact sales) |

## Verdict

Deepgram Voice Agent API sits between the raw OpenAI Realtime API and a fully-configured platform like Bland AI or Retell AI. It is a managed STT, LLM and TTS orchestration layer from the company several rival platforms already plug in as a speech-to-text provider, and LiveKit Agents names Deepgram directly in its own provider list. The genuinely distinguishing fact is deployment. Fully managed cloud, dedicated single-tenant, in-VPC, or fully self-hosted are all real options, backed by SOC 2 Type II certification, signed HIPAA BAAs, GDPR readiness with an EU data residency endpoint, and CCPA and PCI compliance. Its built-in LLM router is also the broadest here. Six named providers spanning 30+ models, including open-weight options via Groq and NVIDIA, versus a fixed model or a narrower bring-your-own list everywhere else, and a fallback-chain feature that lets you list several providers and roll over on failure. The tradeoffs: no ongoing free tier (only a one-time $200 credit), no dashboard or CRM integrations, and the published rates include promotional cuts through September 12, 2026 that buyers signing up after that date will not get.

## How it works

1. Open a WebSocket connection to Deepgram's Voice Agent API and configure the agent: which STT model (Flux or Nova-3), which LLM provider (Deepgram-managed by default, or one of six bring-your-own providers, with an optional fallback chain), which TTS voice (Flux by default, or bring-your-own), and any functions the agent can call mid-conversation.
2. Deepgram's built-in Voice Activity Detection and turn-taking model listen for the caller to finish speaking or interrupt, transcribe in real time, and pass the transcript to the configured LLM for a response.
3. The LLM's response streams back through Deepgram's Flux TTS (or a bring-your-own TTS provider) and plays to the caller, with barge-in handling ready to interrupt playback the instant the caller speaks again.
4. Deploy the same configuration through Deepgram's managed cloud, a dedicated single-tenant instance, your own VPC, or a fully self-hosted install, with EU data residency available on api.eu.deepgram.com.

## Who it's for

- Engineering teams that want a managed speech pipeline (not raw model access) but still need to choose or switch their own LLM provider without being locked to one vendor.
- Regulated enterprises (healthcare, finance, EU public sector) that need a voice-agent vendor able to deploy inside their own VPC, fully on-premise, or with EU data residency.
- Teams already using Deepgram's standalone STT or TTS APIs who want to add LLM orchestration and agent logic without switching speech-infrastructure vendors.

## Strengths and weaknesses

- ✓ The broadest built-in LLM router in this category: six named providers (OpenAI, Anthropic, Google, NVIDIA, Groq, Amazon Bedrock) spanning 30+ models, including open-weight options, versus a fixed model or a narrower bring-your-own list at every other voice-agent listing in this index. Includes a fallback-chain feature for rolling across providers on failure.
- ✓ The only closed-source voice-agent vendor in this category with a genuine self-hosted or VPC-hosted deployment option, alongside fully managed cloud and dedicated single-tenant, plus EU data residency on api.eu.deepgram.com, backed by SOC 2 Type II certification, signed HIPAA BAAs, GDPR, CCPA and PCI compliance.
- ✓ Backed by a well-capitalized, proven speech-infrastructure company: 1,300+ organizations including NASA, Spotify and Citibank, and a $130 million Series C in January 2026 at a $1.3 billion valuation.
- ✓ Mid-conversation function calling is built into the API, not left for a developer to hand-roll, and the LLM router supports an ordered fallback chain across multiple providers for resilience.
- ✗ No ongoing free tier: only a one-time $200 signup credit before usage billing starts, unlike LiveKit Agents' recurring 1,000 free agent-session-minutes per month.
- ✗ No dashboard, call analytics, or CRM or helpdesk integrations bundled in. It is infrastructure a developer builds on top of, not a configured product like Bland AI or Synthflow AI.
- ✗ The published rates include promotional cuts through September 12, 2026. A buyer signing up after that window will pay the post-promo Standard rate of $0.075/min and a $0.0450 per 1k characters Flux TTS fee, materially higher than what the page advertises today.
- ⚠ No dedicated free tier. A one-time $200 credit only, then usage billing from the first paid session.
- ⚠ No bundled dashboard, call analytics, or CRM or helpdesk integrations.
- ⚠ Promotional rates and the free Flux TTS window both expire on 9/12/2026. Post-promo rates run materially higher.

## Key features

- **Flat $4.50/hr headline rate** — Standard tier is a single $4.50/hr figure ($0.075/min) calculated on WebSocket connection time, with promotional cuts through September 12, 2026, and a Custom tier at $0.050/min for fully bring-your-own LLM and TTS.
- **6-provider LLM router** — Route to OpenAI, Anthropic, Google, NVIDIA, Groq or Amazon Bedrock (30+ named models, including open-weight options via Groq and NVIDIA), or let Deepgram manage the LLM by default.
- **Fallback chain for LLM providers** — Pass an ordered array of providers and the Voice Agent rolls to the next one on error or timeout, mixing provider types (for example OpenAI primary, Anthropic fallback) to ride out independent-provider outages.
- **Deployment flexibility** — Fully managed cloud, dedicated single-tenant, in-VPC, or fully self-hosted: a spread no other closed-source vendor in this category offers, with EU data residency on api.eu.deepgram.com.
- **Built-in barge-in and turn-taking** — Voice Activity Detection and a turn-taking model handle interruptions and know when a caller has finished speaking, without a developer building that logic.
- **Mid-conversation function calling** — The agent calls developer-defined functions live during a call and speaks the result back once the client returns it.
- **Flux TTS default during promo window** — Deepgram's Flux TTS is free inside the Voice Agent through September 12, 2026, and is the default unless the developer brings their own TTS.

## Use cases

- **Custom voice agent with LLM flexibility** — An engineering team builds a voice agent that starts on Deepgram's managed LLM for speed, then switches to a specific provider (e.g. Anthropic Claude Sonnet 5) as requirements change, without re-architecting the speech pipeline.
- **Regulated-industry deployment inside a private VPC** — A healthcare or financial-services team deploys the same Voice Agent API configuration inside its own VPC or fully self-hosted, keeping call audio and transcripts inside infrastructure it controls, backed by Deepgram's HIPAA BAA.
- **EU data residency on the same product** — An EU public-sector or healthcare team points the agent at api.eu.deepgram.com and gets data residency plus a self-hosted or single-tenant option, without adopting a separate EU-only vendor.
- **Adding agent logic to an existing Deepgram STT or TTS integration** — A team already using Deepgram's standalone transcription or speech-synthesis APIs adds the Voice Agent API's orchestration layer to turn a transcription pipeline into a full conversational agent, without switching speech vendors.

## Integrations

LLM providers: OpenAI, Anthropic, Google (AI Studio and Gemini Enterprise Agent), NVIDIA, Groq, Amazon Bedrock (bring-your-own or Deepgram-managed) · Telephony: WebSocket-based; SIP or PSTN connectivity via a telephony provider · EU data residency endpoint at api.eu.deepgram.com · No native CRM, helpdesk or dashboard integrations. A developer wires session events and function calls into their own systems.

## Sources

- https://deepgram.com/product/voice-agent-api
- https://deepgram.com/pricing
- https://developers.deepgram.com/docs/voice-agent-llm-models
- https://developers.deepgram.com/docs/voice-agent
- https://techstartups.com/2026/01/13/voice-ai-startup-deepgram-raises-130m-at-1-3-billion-valuation-to-fuel-global-expansion/
- https://www.accountablehq.com/post/is-deepgram-hipaa-compliant-baas-phi-and-security-explained
