AI Voice Agents: Vapi vs Retell AI vs ElevenLabs
Every voice agent we index, compared by the layer of the stack you own: raw model, developer platform, no-code builder, or managed incumbent.
On this page
Scope · 13 topics
- Voice agents
- Decision guide
- OpenAI Realtime API
- Vapi
- Retell AI
- Bland AI
- ElevenLabs Conversational AI
- Synthflow AI
- LiveKit Agents
- PolyAI
- Deepgram Voice Agent API
- Cartesia
- VoiceAgent
Eleven names dominate the shortlist whenever a team starts building an AI agent that answers or makes phone calls: OpenAI Realtime API, Vapi, Retell AI, Bland AI, ElevenLabs Conversational AI, Synthflow AI, LiveKit Agents, PolyAI, Deepgram Voice Agent API, Cartesia and VoiceAgent. Eight of them handle real-time speech, turn-taking and tool-calling on a live call as a platform you configure. The other three, OpenAI's Realtime API, Deepgram's Voice Agent API and Cartesia, aren't platforms in the same sense: OpenAI's is the raw speech-to-speech model some of the others let you plug in directly, Deepgram's is the speech-infrastructure company's own orchestration layer around STT/LLM/TTS components some of the others also plug in as a provider, and Cartesia is a related but distinct third shape: its own Sonic/Ink speech models already power part of this category from behind the scenes (Vapi, Retell AI and LiveKit Agents are all named Cartesia customers), with Line, an open-source agent SDK, built on top. Read their own researched listings side by side and the real split isn't "which one sounds best," it's which layer of the voice stack you want to own, whether you want to rent that stack at all, and who on your team is meant to build the agent; plus, for one name on this list, which specific languages the agent needs to speak. The OpenAI Realtime API sits at one extreme: no platform, no dashboard, just OpenAI's own model billed purely on audio/text tokens: the building block LiveKit Agents' own plugin ecosystem lists by name. Deepgram Voice Agent API sits close to it but not identical: a managed STT+LLM+TTS pipeline with a genuinely broad built-in LLM router (six providers), deployable fully managed, VPC-hosted, or fully self-hosted, the only entry here with that deployment spread. Cartesia sits near both: an infrastructure supplier like Deepgram, but its Line SDK is the most open agent layer on this page (Apache 2.0, 100+ LLM providers via LiteLLM). The tradeoff is that only the underlying speech models, not Line's own runtime, can be self-hosted, and only on an Enterprise contract. Bland AI runs its own platform infrastructure end to end and sells that as compliance and predictability; Vapi and Retell AI both hand you a bring-your-own-model, bring-your-own-voice builder and compete on developer control; ElevenLabs starts from the opposite end (its own best-in-class text-to-speech and cloning models) and wraps an agent layer around them; Synthflow AI is the one platform built for the ops/CX team rather than the engineer, with a visual Flow Designer replacing an API as the primary interface; LiveKit Agents skips the "which vendor" question entirely and hands you an open-source framework to self-host; PolyAI is the pre-LLM-era enterprise incumbent (nine years of production history and 200+ named enterprise customers) reached entirely through a sales conversation, the same access model it ran for most of its life (a self-serve on-ramp it briefly opened in May 2026 is gone from its own site as of an August 2026 re-check); and VoiceAgent is the one name here that competes on language rather than architecture, built specifically for calls that move between Yoruba, Igbo, Hausa, Nigerian Pidgin and Nigerian English inside a single conversation, self-serve from $12.50/month. Which of those starting points you want usually settles the shortlist before a single feature comparison matters.
The eleven, at a glance
Every figure below is pulled from each tool's own researched listing, last verified 13 July – 23 August 2026.
| Agent | Stack ownership | Headline pricing | Real all-in cost | Best known for |
|---|---|---|---|---|
| OpenAI Realtime API | Not a platform at all: the raw model, run by OpenAI | $32/$64 per 1M audio tokens in/out (gpt-realtime-2.1); $10/$20 on the mini | ~$0.06-0.11/min with caching (independently reported); no platform fee, but no dashboard either | Native SIP telephony; the model layer several other platforms here plug in |
| Bland AI | Self-hosted, owns the full stack | $0.14/min (Start) down to $0.11/min (Scale), plus tier platform fees | All-inclusive: no token or model pass-throughs on top | Compliance: SOC 2, HIPAA, PCI DSS, self-host/on-prem, US/EU/APAC residency |
| Vapi | Bring your own STT/LLM/TTS; Vapi orchestrates | $0.05/min orchestration fee | ~$0.13–$0.33/min once model, voice and telephony are added | Cleanest developer API; model- and voice-agnostic by design |
| Retell AI | Bring your own LLM/voice/telephony; Retell orchestrates | $0.07–$0.31/min pay-as-you-go | ~$0.13–$0.20+/min fully specced (voice infra + TTS + telephony + LLM) | Voice quality and latency most-praised in user reviews (G2) |
| ElevenLabs Conversational AI | ElevenLabs' own TTS/cloning models; LLM is bring-your-own-or-hosted | $0/$6/$22/$99/$299/$990/mo tiers (minute pools + concurrency caps) | Tier fee + per-minute overage + separate LLM/telephony costs | 10,000+ voices, instant cloning, 70+ languages; $11B-valued |
| Synthflow AI | No-code Flow Designer; OpenAI models only, bring-your-own or Synthflow-managed telephony | $0.09/min Voice Engine fee | ~$0.11–$0.16/min once an LLM and telephony are added; no self-serve monthly plan | No-code visual builder + a named BELL deployment methodology for ops/CX teams |
| LiveKit Agents | Open-source (Apache 2.0); self-host entirely or use LiveKit Cloud | Free up to 1,000 agent-session min/mo, then $50/mo (5k min) or $500/mo (50k min) | Base meter + a separate LLM/STT/TTS bill (LiveKit Inference credits or your own provider keys) | Powers OpenAI's ChatGPT Advanced Voice Mode; 11.5k+ GitHub stars |
| PolyAI | Proprietary Raven model, or bring-your-own reasoning model (GPT-5/Claude/Gemini) since May 2026 | Enterprise: custom per-minute, sales-led (a self-serve trial it ran May–~Aug 2026 is no longer live) | No public rate card; bundled per-minute rate; third-party estimates put contracts from ~$150,000/year | 200+ enterprise customers, 75+ languages, 25+ countries; $750M valuation; nine years of production history |
| Deepgram Voice Agent API | Deepgram's own managed STT+LLM+TTS orchestration; BYO LLM/TTS to cut cost; fully managed, VPC-hosted or self-hosted | $0.075/min (Standard) down to $0.050/min (Custom, BYO LLM+TTS) | $200 one-time trial credit, then usage-based from the first paid session; up to 20% off on the annual Growth Plan | Broadest built-in LLM router (OpenAI, Anthropic, Google, NVIDIA, Groq, Bedrock); the only self-hostable closed-source option here |
| Cartesia | Cartesia's own Sonic/Ink speech models; Line (open-source SDK or no-code builder) orchestrates, BYO any of 100+ LLMs | $0/$5/$49/$299/mo tiers (model credits) + flat $0.06/min agent minutes on every tier | Credit pool covers Sonic/Ink usage; agent minutes bill separately at $0.06/min ($0.014/min telephony) regardless of plan | Speech models several rival platforms here already license (Vapi, Retell AI, LiveKit Agents are Cartesia customers); most open agent SDK on this page |
| VoiceAgent | Closed-source hosted platform, self-serve | $12.50/mo Starter (150 min included, $0.12/min overage) | Overage-based once included minutes run out; no separate LLM/telephony pass-through to price | Only entry trained for Yoruba, Igbo, Hausa and Nigerian Pidgin, including mid-call code-switching |
The decision: which layer do you want to own?
All ten carry a full researched verdict on their own listing, including who each one is explicitly not for. Read together, the honest decision tree runs on stack ownership and who's building it first, features second:
- You want to go straight to the model with no platform, no dashboard and pure usage-based token pricing, and you're comfortable being single-vendor on OpenAI and writing everything else yourself → OpenAI Realtime API. It's OpenAI's own speech-to-speech model, billed purely on audio/text tokens ($32/$64 per 1M in/out on the flagship, a third of that on the mini), with native SIP telephony built in. Watch out for: no dashboard, no CRM integrations, 60-minute session caps, and no free tier at all: session management, tool-calling and observability are all code you write.
- You're a regulated enterprise and want one vendor to own the entire hosted stack, with a predictable all-in rate and no model pass-throughs → Bland AI. Its Conversational Pathways give deterministic, auditable call flows, and its per-minute price already bundles the LLM, transcription and voice with SOC 2, HIPAA and PCI DSS credentials behind it. Watch out for: independent reviews cite real-world latency nearer ~800ms than marketed, support that leans on public Discord at lower tiers, and full language breadth is effectively gated to enterprise deals.
- You have engineers who want maximum control and the cheapest possible orchestration layer, and you're comfortable assembling and pricing the rest of the stack yourself → Vapi. It's model- and voice-agnostic by design, so you can chase the best model or cheapest voice without rebuilding the agent, and its API is widely cited as the cleanest of the category. Watch out for: the $0.05/min headline is only the orchestration fee: budget for $0.13–$0.33/min once a real model, voice and telephony are stacked on top.
- Voice quality and latency on the actual call matter more to you than headline price, and you want built-in extras like a streaming knowledge base and pre-launch simulation testing → Retell AI. Users consistently praise its natural-sounding calls and ~600ms turn-taking, and it ships an auto-syncing RAG knowledge base, batch outbound calling and an AI QA layer that simulates conversations before they go live. Watch out for: the same component-pricing pattern as Vapi (the $0.07/min headline realistically lands at $0.13–$0.20+/min once fully specced), plus some reviews flag a limited visual builder and slow support.
- Voice quality and multilingual reach are the deciding factor, not the orchestration platform underneath → ElevenLabs Conversational AI. It starts from the industry's most-cited text-to-speech and instant-cloning models, natively speaks 70+ languages, and lets you reason with an ElevenLabs-hosted model for the lowest latency or bring your own LLM key, including open-source models. Watch out for: it covers the voice/LLM-orchestration layer only (you still assemble telephony, CRM/helpdesk workflow and observability yourself), and its free tier (15 minutes/month) is a demo, not a production trial.
- You don't have engineers to spare and want an ops/CX team to design and launch the agent visually, with a documented compliance set already in place → Synthflow AI. Its Flow Designer is genuinely no-code, its BELL methodology (Build, Evaluate, Launch, Learn) structures the rollout, and it lists SOC 2, HIPAA, GDPR, PCI DSS and ISO 27001 together. Watch out for: the same component-pricing pattern as Vapi and Retell AI (a real configuration lands near $0.11–$0.16/min), no self-serve monthly plan (only metered usage or a ~$30,000/year Enterprise contract), and LLM choice limited to OpenAI models.
- You want to own the stack yourself (self-hosted, no vendor lock-in, full code visibility) and have the engineering time to assemble it → LiveKit Agents. It's the only genuinely open-source (Apache 2.0) option here, proven at serious scale (it powers OpenAI's ChatGPT Advanced Voice Mode, plus xAI, Salesforce and Tesla), with a real free tier (1,000 agent-session minutes/month, no card). Watch out for: it's a framework, not a dashboard: you write and deploy code, and the free tier's small inference credit means real usage still means paying LiveKit or a model provider on top.
- You want a proven, enterprise-scale vendor with nine years of production history and hundreds of named customers, and you're prepared to start with a sales conversation rather than a self-serve signup → PolyAI. It's the oldest name in the category by years, with 200+ enterprise customers and thousands of live deployments across 75+ languages, reachable via a no-code builder or a developer-focused ADK once an engagement begins. Watch out for: no public rate card at all (third-party estimates put contracts from ~$150,000/year), and a self-serve on-ramp it briefly ran from May 2026 is gone from its own site as of an August 2026 re-check.
- You want the speech-infrastructure company's own product, the broadest built-in choice of LLM, and a real option to self-host or run in your own VPC instead of trusting a vendor's cloud alone → Deepgram Voice Agent API. It's a managed STT+LLM+TTS pipeline that routes across six LLM providers (OpenAI, Anthropic, Google, NVIDIA, Groq, Amazon Bedrock) out of the box and, uniquely among this page's closed-source vendors, deploys fully managed, dedicated single-tenant, VPC-hosted, or fully self-hosted, backed by SOC 2 Type II and HIPAA BAAs. Watch out for: no ongoing free tier (a one-time $200 signup credit only), no dashboard or CRM integrations (it's infrastructure, not a configured product), and per-minute pricing that only drops to its cheapest tier once you bring your own LLM and TTS.
- You want speech-model infrastructure that already powers part of this category, paired with the most open-source agent layer on this page and a genuine $0/month tier → Cartesia. Its Sonic/Ink speech models are already used by Vapi, Retell AI and LiveKit Agents as a component, and Line, its agent SDK, is Apache 2.0 and connects to 100+ LLM providers via LiteLLM: no other entry on this page ships an open-source agent layer built on top of infrastructure this widely used. Watch out for: unlike Deepgram, only the Sonic/Ink models, not Line's voice-agent runtime itself, can be self-hosted, and only on an Enterprise contract; and pricing is a two-meter system (a model-credit subscription plus a separate agent-minute balance) rather than one blended rate.
- Your calls need to run in Yoruba, Igbo, Hausa or Nigerian Pidgin, sometimes switching between them and English mid-call → VoiceAgent. It's the only platform on this page built around that specific language set rather than general multilingual breadth, and self-serve pricing starts at $12.50/month for 150 included minutes. Watch out for: none of the SOC 2, HIPAA or PCI DSS credentials the bigger vendors here list are part of its public profile, and its signup credit is advertised without a stated dollar figure.
OpenAI Realtime API: not a platform at all, the model underneath several of them
Every other entry on this page is a platform you configure. Even LiveKit Agents, the open-source one, ships session management, turn detection and telephony plumbing you build on top of. The OpenAI Realtime API skips that layer entirely: it's OpenAI's own speech-to-speech model, and you talk to it directly over WebRTC, WebSocket or native SIP, with no dashboard, no bundled telephony vendor and no orchestration code included. That's not a gap in an otherwise-complete product: it's the whole pitch. Pricing follows the same logic: not a per-minute rate at all, but raw audio/text token billing ($32/$64 per 1M tokens in/out on the flagship gpt-realtime-2.1, $10/$20 on the smaller gpt-realtime-2.1-mini), which independently reported production usage translates to roughly $0.06–$0.11/min once prompt caching is working. Native SIP telephony means real inbound/outbound phone calls connect without a separate vendor like Twilio for the connection itself, and OpenAI's own API Platform carries SOC 2 Type 2 and ISO/IEC 27001-family certification. The catches are exactly what you'd expect from going straight to the model: sessions cap at 60 minutes, there's no free tier or trial credit specific to the Realtime API, and every other platform on this page's built-in features (a knowledge base, call analytics, CRM sync, a hosted dashboard) are code a developer writes themselves. It's also the one entry here several of the others could plug into: LiveKit Agents' own plugin ecosystem lists OpenAI by name as a pluggable speech-to-speech provider, and independent "hybrid stack" write-ups describe pairing OpenAI's reasoning with ElevenLabs' voice and Twilio's transport rather than using any one vendor's full stack.
Deepgram Voice Agent API: the speech-infrastructure company's own stack, not a raw model
Deepgram is the company several other entries on this page already plug in as a component: LiveKit Agents' own provider-agnostic plugin list names Deepgram by name for speech-to-text, alongside AssemblyAI, OpenAI, Anthropic and ElevenLabs. The Voice Agent API is Deepgram's own answer to "why assemble that pipeline yourself": a managed STT+LLM+TTS orchestration layer with built-in barge-in detection, turn-taking prediction and mid-session function calling, sitting a level above the raw OpenAI Realtime API (a single model call) and a level below a fully-hosted dashboard product like Bland AI. The default configuration routes through Deepgram's own STT and TTS with a managed LLM; bringing your own LLM and/or TTS provider both unlocks a genuine choice of six named LLM providers (OpenAI, Anthropic, Google, NVIDIA, Groq and Amazon Bedrock, spanning 30+ specific models, including open-weight options via Groq and NVIDIA) and lowers the per-minute rate, from $0.075/min (Standard, Deepgram's own STT+TTS) down to $0.050/min (Custom, fully bring-your-own on both LLM and TTS). The genuinely distinguishing fact for this category is deployment: fully managed, dedicated single-tenant, VPC-hosted, or fully self-hosted are all real options, backed by SOC 2 Type II certification and signed HIPAA BAAs. No other closed-source vendor on this page offers a self-hosted or VPC path at all, only Bland AI (also proprietary) and the open-source LiveKit Agents come close, and neither matches Deepgram's specific combination of closed engine plus deployment choice. The catches: like OpenAI's Realtime API, there's no ongoing free tier, only a one-time $200 signup credit before usage billing starts; there's no dashboard, call analytics or CRM integration bundled in: it's infrastructure a developer builds on top of, the same shape as Vapi and Retell AI rather than a configured product like Bland AI or Synthflow AI; and the published Standard/Advanced pricing tiers correspond to different underlying model tiers Deepgram's own docs don't fully break down on the public pricing page.
Cartesia: the other speech-infra company, but the most open agent SDK here
Cartesia and Deepgram share a shape (both are speech-infrastructure companies that built an agent layer on top of models several rival platforms already license as a component), but they diverge sharply on openness and deployment. Cartesia's own customer page names Vapi, Retell AI and LiveKit Agents directly as users of its Sonic (TTS) and Ink (STT) models, the same kind of provider-ecosystem credibility Deepgram and OpenAI carry in this category. Line, Cartesia's agent layer built on top, is genuinely Apache 2.0: a real, functional SDK (not a thin wrapper) supporting multi-agent handoffs, tool calling and a no-code Agent Builder for prototyping, connecting to more than 100 LLM providers via LiteLLM rather than locking a buyer into one reasoning model. Pricing follows a two-meter pattern distinct from every other entry here: a monthly credit subscription ($0 to $299) covers Sonic/Ink usage, while agent minutes bill separately at a flat $0.06/min ($0.014/min telephony) that doesn't change across tiers, and the $0/month Free tier is a genuine ongoing plan, not a time-limited trial. The deployment story runs opposite to Deepgram's, though: only the underlying Sonic and Ink models can be self-hosted, on an Enterprise contract into a customer's own cloud or on-premise environment. Line's own voice-agent runtime always stays on Cartesia's managed infrastructure, at every tier, with no self-hosted or VPC path at all. And unlike Retell AI, Vapi or Bland AI, Cartesia hasn't published an independent review-platform rating (G2, Capterra, Trustpilot) as of this guide's research pass, despite real production customers and a $191M raise across seed, Series A and Series B rounds (most recently $100M in October 2025, led by Kleiner Perkins with NVIDIA joining as a new strategic investor).
LiveKit Agents: a fifth shape entirely, not a rival to the other four
Bland AI, Vapi, Retell AI and ElevenLabs Conversational AI are all closed-source SaaS: you rent an API or a dashboard, and the vendor's infrastructure is the product. LiveKit Agents doesn't compete on that axis at all: it's an Apache 2.0-licensed, open-source framework (11.5k+ GitHub stars) you write Python or TypeScript code against, and you can run the entire stack (speech-to-text, LLM orchestration, text-to-speech, telephony) on your own infrastructure, with no vendor able to change your pricing or shut off your access. LiveKit Cloud exists as an optional managed layer if you'd rather not run servers yourself, but nothing about the framework requires it. That openness is backed by real production weight: LiveKit is the company behind OpenAI's ChatGPT Advanced Voice Mode, and its other named customers include xAI, Salesforce and Tesla. This isn't an unproven side project, it's infrastructure already running at some of the largest voice-AI deployments anywhere. The honest tradeoff is the one every open-source framework carries versus a configured SaaS product: there's no hosted dashboard that gets a non-technical team a working agent in an afternoon, and the free tier's inference credit (~50 minutes of LiveKit-hosted model usage) is small enough that any real deployment needs a plan for ongoing LLM/STT/TTS costs on top of LiveKit's own agent-session-minute meter.
Synthflow AI: the no-code shape, not a fourth developer-first rival
Vapi, Retell AI and ElevenLabs Conversational AI are all built for an engineer to configure through an API. Synthflow AI is built the opposite way: its primary interface is a visual Flow Designer, and its named BELL Framework (Build, Evaluate, Launch, Learn) turns rollout into a structured process rather than something each team invents for itself, a real difference in who can operate the platform, not just a styling choice on top of the same product. That's backed by real scale for the funding stage: a $20M Series A led by Accel in June 2025 ($30M raised total) and more than 45 million calls processed since its 2023 Berlin founding, plus an unusually complete compliance set for a company this young: SOC 2, HIPAA, GDPR, PCI DSS and ISO 27001 all listed together, with EU/US regional hosting. The honest catches: pricing follows the same three-component pattern as Vapi and Retell AI (a $0.09/min Voice Engine fee, then a separate LLM fee and telephony fee stacked on top, landing near $0.11–$0.16/min in a working configuration), there's no self-serve monthly subscription (only metered pay-as-you-go or a custom Enterprise contract starting around $30,000/year), and its published LLM choice is limited to OpenAI models, where Vapi and Retell AI both advertise full bring-your-own-model flexibility.
Vapi vs Retell AI: the closest real head-to-head
Of the four hosted platforms, Vapi and Retell AI are the actual rivals: both are developer-first, bring-your-own-model platforms with no self-hosted or fully-managed alternative, competing for the same technical-team budget. The difference is what's built in versus what you assemble yourself. Vapi's pitch is control and provider-agnosticism taken furthest: pick any STT, LLM and TTS, and coordinate multiple specialised agents ("squads") that hand off mid-call. Retell AI leans further into shipped tooling on top of the same bring-your-own-stack idea: an auto-syncing knowledge base, batch outbound calling and built-in simulation testing that Vapi doesn't list as native features, and user reviews credit it with the more natural-sounding voice output of the two. Neither publishes a true all-in price: both headline a base orchestration or pay-as-you-go rate that reviewers and the vendors' own docs confirm climbs several times higher once a real model, voice and telephony are added, so budgeting either one requires pricing the full stack, not the homepage number.
PolyAI: the pre-LLM enterprise incumbent, sales-led once again
PolyAI predates every other name on this list by years: founded in London in 2017, when Vapi, Retell AI, Bland AI, ElevenLabs Conversational AI and LiveKit Agents didn't yet exist. That history shows up as a scale gap: 200+ named enterprise customers, thousands of live deployments across 75+ languages in 25+ countries, and an $86M Series D (December 2025) that pushed its valuation to $750M, with NVIDIA's NVentures among the backers. Its Raven dialog model, trained on over a billion enterprise conversations, now also runs alongside third-party reasoning models (GPT-5, Claude, Gemini) since a May 2026 platform update. That same platform update briefly ended PolyAI's sales-led-only era: for a few months, a team could sign up with just an email address and get two months of free access, either through a no-code Poly Agent Builder or a developer-focused Agent Development Kit (ADK). Re-verified live in August 2026, that self-serve door is closed again: PolyAI's own pricing and marketing pages no longer mention a free period or email signup, and independent 2026 pricing reviews now describe it as sales-led only, with no free tier or trial. So the honest current state matches PolyAI's first eight years: both build tools, and every deployment (including PG&E's, 16 million calls a year, with tens of thousands more during emergency spikes), run through a custom-quoted enterprise engagement, with no public rate card at all. A vendor-commissioned but methodologically transparent Forrester study modelling 391% three-year ROI adds real, checkable numbers that most of the newer entrants here don't have a decade of customer relationships to produce.
VoiceAgent: the only platform here built for West African languages, not just more of them
Every other name on this page treats language breadth as a checkbox: ElevenLabs advertises 70+ languages, PolyAI claims 75+, and the OpenAI Realtime API and Deepgram's LLM router both inherit whatever languages their underlying models support. None of them are built around one specific regional language set the way VoiceAgent is. It runs calls in Yoruba, Igbo, Hausa, Nigerian Pidgin and Nigerian English, and (the harder capability the generalist platforms don't claim at all) switches between them inside a single call the way a bilingual Nigerian call-center agent would, rather than requiring a caller to pick one language up front. That's a different product decision than the other ten make: Vapi, Retell AI and Bland AI all assume a team supplies its own language coverage through whichever LLM and TTS provider it plugs in, and none of the providers those three would plug in (OpenAI, ElevenLabs, Deepgram) publish trained code-switching for this specific language set. Pricing follows the self-serve subscription shape of ElevenLabs or Cartesia rather than the developer per-minute meter of Vapi or Retell AI: a Starter plan at $12.50/month includes 150 minutes with $0.12/min overage, stepping up to Growth at $40.80/month and Business at $165.80/month, with an Enterprise tier custom-quoted above that and an undisclosed signup credit applied at signup. The catch is the flip side of building for one regional market first: none of the SOC 2, HIPAA or PCI DSS credentials that Bland AI, Deepgram, Cartesia and Synthflow AI all list publicly are part of VoiceAgent's own site, and unlike Deepgram's explicit $200 signup credit, VoiceAgent doesn't state the size of its own.
Ten platforms publish a usage-based rate; PolyAI publishes none at all
Vapi, Retell AI, Bland AI, ElevenLabs Conversational AI, Synthflow AI, LiveKit Agents, Deepgram Voice Agent API, Cartesia and VoiceAgent all bill by usage (per-minute, a metered credit pool, or both), not a flat per-seat licence, which fits a category built to replace or augment human phone work rather than seats at a desk. PolyAI's enterprise tier follows the same logic in principle (a custom, sales-led per-minute rate bundling tuning, maintenance and support), but breaks the pattern in a different way: it's the only one of the eleven with no public rate card at all, so there's no headline number to even partially compare: third-party estimates put typical contracts starting around $150,000/year. The OpenAI Realtime API breaks the pattern a third way: it bills by usage too, but in raw model tokens rather than a per-minute rate, with no platform fee wrapped around it at all. None of the other ten publishes a genuine all-in per-minute number either: Bland AI comes closest by bundling the LLM and voice into one rate, yet even its price varies by tier and platform fee; Vapi, Retell AI, ElevenLabs, Synthflow AI, LiveKit Agents, Deepgram and Cartesia all headline a partial figure (an orchestration fee, a pay-as-you-go range, a subscription-plus-overage tier, a base voice-engine or agent-session-minute meter, or a two-meter credit-plus-agent-minute system) that only becomes a real number once you've picked a model, voice and telephony provider and stacked their costs on top. VoiceAgent's $12.50/month Starter is the closest anything here gets to a single all-in number for a small deployment (150 minutes for one flat fee, no separate model or telephony bill to add), but it's still a subscription-plus-overage tier rather than a blended per-minute rate, and it only holds for the first 150 minutes a month. Deepgram's own per-minute rate is close, but its cheapest published tier ($0.050/min) only applies once you've brought your own LLM and TTS, the same "headline isn't the real number" pattern as Vapi and Retell AI; Cartesia's $0.06/min agent rate is flat across tiers, but it's layered on top of a separate model-credit subscription most other entries here don't have. The OpenAI Realtime API's token price is the most transparent number on this page, but it still needs converting into a per-minute estimate to compare against the rest (roughly $0.06–$0.11/min on the flagship model with caching, independently reported) since nothing about the raw rate says "per minute" on its own. Any per-minute figure you see quoted for this category, on these listings or anywhere else, is a starting point for a spreadsheet, not a number to budget against directly.
FAQ
Which of these eleven owns its own infrastructure end to end?
Five do, in different ways: Bland AI runs its own model stack end to end as a hosted service and charges one all-inclusive rate, OpenAI Realtime API is the model itself, run directly by OpenAI, with no separate platform layer wrapped around it, Deepgram Voice Agent API is the speech-infrastructure vendor's own orchestration layer, deployable fully managed, VPC-hosted, or fully self-hosted, Cartesia owns its Sonic/Ink speech models end to end (several rival platforms on this page plug them in as a component) though its Line agent runtime always stays on Cartesia's own infrastructure rather than the customer's, and LiveKit Agents is open-source (Apache 2.0) and fully self-hostable: you can run the entire stack on your own infrastructure, or use LiveKit's managed Cloud if you'd rather not. Vapi, Retell AI, ElevenLabs, Synthflow AI, PolyAI and VoiceAgent are all closed-source hosted platforms you can only rent, though PolyAI, like the others, now lets you bring a third-party reasoning model (GPT-5, Claude or Gemini) rather than only its own proprietary Raven model, reachable once a sales engagement is underway; Synthflow AI's own model choice is limited to OpenAI models.
Which of these agents supports Nigerian or West African languages?
Only VoiceAgent is built specifically for that: Yoruba, Igbo, Hausa, Nigerian Pidgin and Nigerian English, including switching between them inside a single call. ElevenLabs lists 70+ languages and PolyAI lists 75+, both far broader in raw count, but neither claims trained code-switching for this specific regional set, and neither markets to the Nigerian market the way VoiceAgent does. A team whose calls are mostly English with occasional other-language support is better served by ElevenLabs' or PolyAI's general multilingual breadth; a team whose calls genuinely move between Nigerian languages mid-conversation is the specific case VoiceAgent is built for.
Which is cheapest to start using?
By headline rate, Vapi's $0.05/min orchestration fee is the lowest number on the page. For an actually usable free tier, LiveKit Agents (1,000 agent-session minutes/month, no credit card) is far more generous than ElevenLabs' $0 tier (15 minutes/month, a demo allowance) or Cartesia's $0 tier (20,000 model credits plus a $1 prepaid agent balance, worth roughly 16 agent minutes); and Synthflow AI has no free tier at all beyond building/testing in the dashboard. VoiceAgent has no free tier either, but its $12.50/month Starter (150 minutes included) is the cheapest fixed monthly commitment on the page for a team that wants a predictable bill rather than a per-minute meter. OpenAI Realtime API, Deepgram Voice Agent API and PolyAI have no ongoing free tier either: OpenAI bills purely on tokens from the first session, Deepgram gives a one-time $200 signup credit before usage billing starts, and PolyAI's own brief May 2026 self-serve trial is gone from its site as of an August 2026 re-check, leaving a sales conversation as the only way in. Headline price is not real cost for any of the eleven: Vapi, Retell AI, Synthflow AI, LiveKit, Deepgram and Cartesia all bill separately for at least some of the model/voice/telephony components a full deployment needs, Bland AI's and PolyAI's headline numbers both vary by tier or contract, VoiceAgent's $12.50/month covers only the first 150 minutes before per-minute overage applies, and ElevenLabs' free tier is a demo allowance, not a production budget.
What's the real difference between Vapi and Retell AI?
Mostly what ships built in versus what you assemble, not raw capability; see the head-to-head section above. Vapi pushes provider-agnosticism and multi-agent "squads" furthest; Retell AI adds a streaming knowledge base, batch calling and simulation testing on top of the same bring-your-own-stack model, and is more often praised for voice quality in reviews.
Is ElevenLabs Conversational AI a full voice-agent platform or just a TTS add-on?
It's a full agent layer (real-time transcription, an LLM turn and tool-calling, not just text-to-speech), but it covers the voice/orchestration layer only. Telephony beyond its native/Twilio connectors, CRM/helpdesk workflow and observability are left for you to assemble, the same limitation reviewers flag on the bring-your-own-stack platforms.
Is Synthflow AI a good fit for a team without engineers?
Yes: that's its explicit positioning. The Flow Designer is a visual, no-code builder, and the BELL Framework (Build, Evaluate, Launch, Learn) is a documented process for taking a flow from idea to live agent without writing orchestration code. It still ships a full REST API for teams that later want programmatic control, but the primary interface is the visual builder, unlike Vapi, Retell AI and ElevenLabs, which all assume a developer configuring an API from the start.
Is LiveKit Agents a turnkey product or something else entirely?
Something else entirely: it's an open-source framework and SDK, not a configured product. You write agent code (Python or TypeScript) that wires together a speech-to-text, LLM and text-to-speech pipeline, then deploy it yourself or to LiveKit Cloud. That's real engineering work the other platforms mostly hide behind a dashboard, but it also means no vendor can lock you in, raise prices unilaterally, or shut off access to your own infrastructure. The OpenAI Realtime API takes that even further: it isn't a framework either, just the model, with no session/turn-detection plumbing included at all; LiveKit Agents can plug that same model in underneath its own framework.
Which of these eleven is built for regulated industries?
Bland AI is built for this explicitly: it's the only hosted platform marketed at healthcare, finance and insurance with SOC 2, HIPAA and PCI DSS compliance plus self-hosted/on-premises deployment options built in from the start. Deepgram Voice Agent API is a close second: SOC 2 Type II certified with signed HIPAA BAAs, plus the same self-hosted/VPC-hosted deployment flexibility as Bland AI, from a company (Deepgram) whose speech infrastructure already underpins other vendors in this category. Cartesia carries SOC 2 Type II, PCI Level 1, GDPR and HIPAA-eligible compliance on its public API, with a signed BAA available on Enterprise contracts, though, unlike Bland AI and Deepgram, its Line agent runtime itself can't be self-hosted, only the underlying speech models can, and only on that same Enterprise contract. PolyAI has the deepest track record actually running in regulated environments: named customers include a major US utility (PG&E) and a European bank (UniCredit), with SOC 2, HIPAA, GDPR and PCI DSS all shipping on the platform, though every deployment there, regulated-scale or not, now runs through PolyAI's sales-led enterprise path (a self-serve tier it briefly ran from May 2026 is gone from its own site as of an August 2026 re-check). Synthflow AI lists the broadest single compliance set on paper (SOC 2, HIPAA, GDPR, PCI DSS and ISO 27001 together, plus EU/US regional hosting), aimed at BPO, healthcare and financial-services buyers specifically, though it is the newest and least enterprise-proven of the group. OpenAI Realtime API isn't marketed at regulated industries specifically, but OpenAI's own API Platform carries SOC 2 Type 2 and ISO/IEC 27001-family certification a deploying team can build on. LiveKit Agents is the other option worth evaluating: as a self-hosted open-source framework it gives a team full control over where data lives, though compliance certification and audit work then become the deploying team's own responsibility rather than a vendor-provided package. VoiceAgent is the one entry here that doesn't publish any compliance certifications at all: its differentiation is language coverage for the Nigerian market, not a regulated-industry credential set, so a team in a regulated sector would need to vet that separately before relying on it.
How is PolyAI different from the other ten voice agents?
PolyAI is the oldest name in the category by years (founded 2017, a Cambridge ML spinout) and the only one built around a large, named enterprise customer base (200+ customers) rather than a developer-first orchestration layer, a no-code builder, or a raw model API. It briefly opened a free-for-two-months self-serve tier from a May 2026 platform update, but as of an August 2026 re-check that tier is gone from its own site. It is the only one of the eleven with no self-serve access and no public price at all, unlike the fully self-serve, always-on pricing Vapi, Retell AI, ElevenLabs, Synthflow AI, LiveKit Agents, Deepgram Voice Agent API, Cartesia, VoiceAgent and the OpenAI Realtime API all publish.
Every agent above carries its own full researched listing: pricing tiers, sourced pros and cons, and a committed verdict on who it's for and who it isn't. Start with the voice agents category for the full set.
Vapi
Captured · theagentsindex.com Retell AI
Captured · theagentsindex.com Bland
Captured · theagentsindex.com ElevenLabs Agents
Captured · theagentsindex.com Synthflow
Captured · theagentsindex.com LiveKit Agents
Captured · theagentsindex.com PolyAI
Captured · theagentsindex.com OpenAI Realtime API
Captured · theagentsindex.com Deepgram Voice Agent API
Captured · theagentsindex.com Cartesia
Captured · theagentsindex.com VoiceAgents
Captured · theagentsindex.com SkipCalls
Captured · theagentsindex.com
Frequently asked questions
Which of these voice agents owns its own infrastructure end to end?
Which of these agents supports Nigerian or West African languages?
Which is cheapest to start using?
What's the real difference between Vapi and Retell AI?
Is ElevenLabs Conversational AI a full voice-agent platform or just a TTS add-on?
Is Synthflow AI a good fit for a team without engineers?
Is LiveKit Agents a turnkey product or something else entirely?
Which of these agents is built for regulated industries?
How is PolyAI different from the other ten voice agents?
Advertise here
Reach buyers mid-decision. Reach builders choosing their next agent. Promote your brand with a display placement or bring your listing into focus with Featured.
Explore owner options →Advertise on this page →The digestFree
Which agents actually ship.
What we re-checked, what got added, and one number from the index. Tuesdays.
One-click unsubscribe