Skip to content

The datasheet for every AI agent

Deepgram Voice Agent API

Claim

Summary

Deepgram's Voice Agent API is a single WebSocket endpoint that runs the whole voice loop (transcription, LLM orchestration and speech synthesis) inside one service. Turn-taking, barge-in and function calling are handled by the API rather than by glue code you write.

Best for
Engineering teams that want a managed speech pipeline with real LLM-provider choice, and regulated enterprises that need a self-hosted or VPC deployment option a typical closed-source voice-agent vendor does not offer.
Not for
Teams that want a configured, no-code voice-agent product with a dashboard out of the box. Synthflow AI, Retell AI or Bland AI fit that need better. Deepgram Voice Agent API is infrastructure, not a turnkey product.
From
$0.05/minutemetered on usage
Free tier
Partly$200 signup credit, no credit card required and no time limit — quoted as roughly 43,000 minutes of Nova transcription, spendable on any endpoint including the Voice Agent API. It is a credit grant rather than a free tier: promotional credits expire one year from signup and every call is metered once the balance is gone.
Visit Deepgram Voice Agent API CompareFind alternatives

What it does

Done-for-you service
NoSelf-serve software you build against: you sign up, get an API key and write the client. Enterprise brings an account representative and a contract, not a team that builds and runs the agent for you.A single, unified conversational AI API for building enterprise-ready, cost-effective voice AI agents. Combines the simplicity developers want with…deepgram.com/product/voice-agent-api

How much it does unattended

Runs autonomously
YesThe API runs the whole listen-think-speak loop live, handling interruptions and calling functions to take action without a human in the conversation.Our Voice Agent API enables real-time conversational AI agents that seamlessly handle interruptions, take complex actions, and deliver natural,…deepgram.com/product/voice-agent-api
Multi-agent
PartlyHandoff between specialised agents is a documented pattern with a reference implementation you run yourself: your orchestrator opens a fresh Voice Agent session per agent and summarises context between them. The API itself has no built-in agent-to-agent coordination.This implementation creates separate Voice Agent sessions for each specialized agent.developers.deepgram.com/docs/multi-agen…
Agent permissions
PartlyWhat the agent may do is bounded by the function definitions you send in Settings, and you choose whether each function executes client-side or against an endpoint you host. No approval or confirmation gate is documented: once a function is defined, the model may call it mid-conversation.Function Selection: The LLM identifies a matching function from the definitions you provided in your settings.developers.deepgram.com/docs/voice-agen…

Which models it runs on

Claude
YesAnthropic is a Deepgram-managed provider, so no endpoint of your own is needed; Claude Sonnet models bill at the Advanced tier and Claude Haiku at Standard.For open_ai, anthropic, google, and nvidia, the endpoint field is optional because Deepgram provides managed LLMs for these providers.developers.deepgram.com/docs/voice-agen…
GPT
YesOpenAI is a managed provider; the nano and mini GPT models bill at the Standard tier and the full models at Advanced, which roughly doubles the per-minute rate.For open_ai, anthropic, google, and nvidia, the endpoint field is optional because Deepgram provides managed LLMs for these providers.developers.deepgram.com/docs/voice-agen…
Open models
YesOpen-weight LLMs are listed as supported models: gpt-oss-20b through Groq, which requires you to supply the endpoint, and an NVIDIA Nemotron model that Deepgram manages. Deepgram's own speech and voice weights are not distributed.openai/gpt-oss-20bdevelopers.deepgram.com/docs/voice-agen…
Model choice
YesYou choose the STT model, the LLM provider and model, and the TTS voice, and can swap the LLM or voice mid-session with UpdateThink and UpdateSpeak messages.LLM: Use a Deepgram-managed model, bring your own provider, or point at a custom endpoint. The model can call functions and reach your external…developers.deepgram.com/docs/voice-agen…

Where you use it

In your editor
NoThe agent's audio surfaces are browser, mobile app and telephony. The editor-facing tooling Deepgram ships (the CLI's MCP server, a docs MCP server and agent skills) exists to help you build against the API, not to run an agent inside your IDE.Audio streams in from any source — a browser, a mobile app, or a phone call over your telephony provider.developers.deepgram.com/docs/voice-agen…
On the command line
PartlyThe dg CLI covers transcription, synthesis, text analysis and account management; its own getting-started enumeration and command reference contain no voice-agent command, while the agentic-tools page claims the CLI reaches voice agents. Both statements are recorded in notes.Transcribe files, stream live audio, synthesize speech, analyze text, and manage your Deepgram account — without writing a single line of code.developers.deepgram.com/developer-tools…
In a browser
YesAgents can be configured in the Deepgram Console and referenced by ID, tested interactively in the hosted playground, and embedded in any web page through the Browser Agent SDK's drop-in widget.Configure the agent by referencing one you created in the Deepgram Console, or define the full configuration inlinedevelopers.deepgram.com/docs/browser-ag…

Whose machine it runs on

Self-hosted
PartlySelf-hosted containers are an Enterprise-only path. Deepgram's own self-hosting pages enumerate speech-to-text, text-to-speech and audio intelligence, not voice-agent orchestration, while the Voice Agent product page lists self-hosted as a deployment option. Dedicated single-tenant hosting does explicitly cover agent workloads.we offer self-hosted containers that you can deploy in your own VPC or on-premise hardware (requires NVIDIA GPUs).deepgram.com/pricing
Open source
PartlyThe client SDKs are MIT-licensed on Deepgram's own GitHub, and the browser agent packages ship on npm; the Voice Agent runtime, the STT and TTS models and the self-hosted containers are proprietary and license-server validated.MIT License Copyright (c) 2026 Deepgram. Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated…github.com/deepgram/deepgram-js-sdk/blo…

What it costs to run

How it meters
YesMetered per minute of WebSocket connection time, one rate covering STT, LLM and TTS together; the tier depends on which LLM you pick and whether you bring your own LLM or TTS. Billing is true per-second, not rounded up.The Voice Agent API is billed based on the duration of the conversation. It includes the orchestration of Speech-to-Text (listening), LLM processing…deepgram.com/pricing
Free tier
PartlyA $200 signup credit with no credit card required, spendable on any endpoint including the agent, plus a hosted playground. It is a credit grant, not a free tier: once it is spent, every call is metered.Every new account receives $200 in free credit, which is equivalent to approximately 43,000 minutes (over 700 hours) of transcription using our Nova…deepgram.com/pricing
API access
YesThe product is the API: one WebSocket at /v1/agent/converse, with official SDKs for JavaScript, .NET, Python, Go and Java, all marked GA.You open one WebSocket connection, send audio in, and receive audio out. Deepgram runs the full speech loop — speech-to-text, LLM orchestration, and…developers.deepgram.com/docs/voice-agen…
MCP server
PartlyTwo MCP servers exist, one built into the dg CLI and an HTTP docs server, but the CLI server's published tool list is transcription, synthesis, text analysis, model listing and usage. Nothing exposes running or querying a voice agent over MCP.a CLI with a built-in MCP server, a docs MCP server for querying documentation, and a skills marketplace with structured reference material your…developers.deepgram.com/developer-tools…
Bring your own key
YesBring your own LLM, your own TTS, or both, and the rate drops accordingly; Bedrock and Groq require your own endpoint and credentials, and the browser SDK can point at a proxied or self-hosted WebSocket URL.Easily integrate your own LLM or TTS provider while retaining Deepgram's orchestration, streaming pipeline, and real-time responsiveness.deepgram.com/product/voice-agent-api

Buying it for a team

A company can buy it
YesThree purchase paths: self-serve pay-as-you-go, a Growth commitment bought in the Console from $4k/year, and an Enterprise contract with volume discounts, a BAA and marketplace private offers through AWS or GCP.Yes. If your application processes millions of minutes per month, our Enterprise plan offers significant volume discounts, custom model training, and…deepgram.com/pricing
Seat model
NoNo per-seat charge anywhere: team members are invited into a project and billing is per minute of connection time. The published ceiling is concurrency, not users: up to 45 simultaneous agent WebSocket connections on Pay As You Go and 60 on Growth.Voice Agent API Up to 45 for the WSS API Up to 60 for the WSS APIdeepgram.com/pricing
Pooled budget
YesCredits sit on the project, not the user, and every member and API key in that project draws from the same balance; usage can be attributed per customer or team by tagging keys and requests.Projects are assigned credits, which determine how many transactions can be performed for the associated Project.developers.deepgram.com/guides/deep-div…
Admin controls
PartlyRole-based control exists over the account, not over the agent: owner, admin and member roles plus scoped API keys govern who may read usage, create keys or change billing. What the agent itself may say or call is bounded by your prompt and function definitions, with no admin-side policy engine documented.You control what actions a Team Member can perform by assigning them a Role.developers.deepgram.com/guides/deep-div…
Audit log
PartlyPer-request logs (90 days in console, longer via the Usage API, with per-request USD cost) plus per-session message observability; no administrative-action audit trail is documented.You can obtain up to 90 days of log usage data from Deepgram.developers.deepgram.com/docs/using-logs…
Single sign-on
PartlyThe console accepts Google, GitHub and Azure identity logins alongside email and password; no SAML or OIDC enterprise SSO is documented in the security, compliance or roles pages.Google GitHub Azureconsole.deepgram.com/login

What happens to your code

Opt out of training
YesOpt-out is a per-request query parameter, mip_opt_out=true, and Deepgram states that only data contractually included in the Model Improvement Partnership Program is used for training; opted-out request data is retained only long enough to serve the request.Add mip_opt_out=true as a query parameter of all API requests that you want to be excluded from the Model Improvement Program.developers.deepgram.com/docs/the-deepgr…
Data residency
PartlyEU and AU endpoints both serve the agent WebSocket, and Deepgram runs listen and speak in-region; managed LLM traffic is routed in-region only for OpenAI and Google in the EU, and on the AU endpoint the think step is processed outside Australia. Dedicated adds a customer-chosen AWS region.Deepgram runs listen and speak on Australian infrastructure; the think step runs on the third-party provider's infrastructure, which may process data…developers.deepgram.com/reference/regio…
Getting out
PartlyUsage and per-request cost are exportable through the management API and 90 days of request logs, and unused credits are refundable within 30 days of purchase. Expiry depends on origin: card-purchased credits never expire, promotional credits expire a year after signup and contract credits expire with the contract.Deepgram free promotional credits expire one year from signup. Credits purchased by individuals using a credit card do not expire.developers.deepgram.com/guides/deep-div…
Certifications
YesSOC 2 Type 1 and Type 2, HIPAA with a BAA for Enterprise, GDPR with an EU endpoint, CCPA and PCI with yearly review. ISO 27001 is not claimed anywhere on the site.Deepgram has achieved SOC 2 Type 1 and Type 2 certification. An independent auditor has evaluated the security controls and procedures we use to…developers.deepgram.com/trust-security/…

28 sourced claims on this page. Checked 2026-09-02. How we check.

Alternatives

Where we would send a reader instead, in that listing's own words.

  • Vapi

    Developer platform for building, testing and deploying voice AI agents that make and receive phone calls.

  • LiveKit Agents

    Open-source Python and Node.js framework for voice AI agents, with LiveKit Cloud hosting, telephony and observability.

  • Cartesia

    Real-time speech models and a managed runtime for production phone and web voice agents.

  • OpenAI Realtime API

    OpenAI's speech-to-speech API for low-latency voice agents over WebRTC, WebSocket or SIP.

Filed under

APIAutonomousClaudeEnterpriseGPTMulti-agentOpen modelsSelf-hosted

Advertise here

Reach buyers mid-decision. Reach builders choosing their next agent. Promote your brand with a display placement or bring your listing into focus with Featured.

Explore owner options →Advertise on this page →

The digestFree

Which agents actually ship.

What we re-checked, what got added, and one number from the index. Tuesdays.

147 agents trackedOne-click unsubscribe