Vapi vs Retell AI vs Bland AI vs ElevenLabs vs LiveKit vs PolyAI: How to Choose an AI Voice Agent

All 6 voice agents in our index compared side by side — a self-hosted compliance platform, two developer-first bring-your-own-stack builders, a voice-quality specialist, an open-source framework and a pre-LLM-era enterprise incumbent that just opened a self-serve tier are six genuinely different starting points, and the right pick depends on which layer of the stack you actually want to own.

By The Agents Index Editorial, Research desk · 2026-07-21

Six names dominate the shortlist whenever a team starts building an AI agent that answers or makes phone calls: Vapi, Retell AI, Bland AI, ElevenLabs Conversational AI, LiveKit Agents and PolyAI. All six handle real-time speech, turn-taking and tool-calling on a live call — but read their own researched listings side by side and the real split isn't "which one sounds best," it's which layer of the voice stack you want to own, and whether you want to rent that stack at all. Bland AI runs its own infrastructure end to end and sells that as compliance and predictability; Vapi and Retell AI both hand you a bring-your-own-model, bring-your-own-voice builder and compete on developer control; ElevenLabs starts from the opposite end — its own best-in-class text-to-speech and cloning models — and wraps an agent layer around them; LiveKit Agents skips the "which vendor" question entirely and hands you an open-source framework to self-host; and PolyAI is the pre-LLM-era enterprise incumbent — nine years of production history and 200+ named enterprise customers — that only opened a self-serve on-ramp in May 2026, after years of sales-led-only access. Which of those starting points you want usually settles the shortlist before a single feature comparison matters.

The six, at a glance

Every figure below is pulled from each tool's own researched listing, last verified 13–23 July 2026.

AgentStack ownershipHeadline pricingReal all-in costBest known for
Bland AISelf-hosted, owns the full stack$0.14/min (Start) down to $0.11/min (Scale), plus tier platform feesAll-inclusive — no token or model pass-throughs on topCompliance: SOC 2, HIPAA, PCI DSS, self-host/on-prem, US/EU/APAC residency
VapiBring your own STT/LLM/TTS; Vapi orchestrates$0.05/min orchestration fee~$0.13–$0.33/min once model, voice and telephony are addedCleanest developer API; model- and voice-agnostic by design
Retell AIBring your own LLM/voice/telephony; Retell orchestrates$0.07–$0.31/min pay-as-you-go~$0.13–$0.20+/min fully specced (voice infra + TTS + telephony + LLM)Voice quality and latency most-praised in user reviews (G2)
ElevenLabs Conversational AIElevenLabs' own TTS/cloning models; LLM is bring-your-own-or-hosted$0/$6/$22/$99/$299/$990/mo tiers (minute pools + concurrency caps)Tier fee + per-minute overage + separate LLM/telephony costs10,000+ voices, instant cloning, 70+ languages; $11B-valued
LiveKit AgentsOpen-source (Apache 2.0); self-host entirely or use LiveKit CloudFree up to 1,000 agent-session min/mo, then $50/mo (5k min) or $500/mo (50k min)Base meter + a separate LLM/STT/TTS bill (LiveKit Inference credits or your own provider keys)Powers OpenAI's ChatGPT Advanced Voice Mode; 11.5k+ GitHub stars
PolyAIProprietary Raven model, or bring-your-own reasoning model (GPT-5/Claude/Gemini) since May 2026Self-serve: free for 2 months (email signup only); Enterprise: custom per-minute, sales-ledSelf-serve: undisclosed after the 2-month trial. Enterprise: bundled per-minute rate, no public rate card200+ enterprise customers, 75+ languages, 25+ countries; $750M valuation; nine years of production history

The decision: which layer do you want to own?

All six carry a full researched verdict on their own listing, including who each one is explicitly not for. Read together, the honest decision tree runs on stack ownership first, features second:

  • You're a regulated enterprise and want one vendor to own the entire hosted stack, with a predictable all-in rate and no model pass-throughsBland AI. Its Conversational Pathways give deterministic, auditable call flows, and its per-minute price already bundles the LLM, transcription and voice with SOC 2, HIPAA and PCI DSS credentials behind it. Watch out for: independent reviews cite real-world latency nearer ~800ms than marketed, support that leans on public Discord at lower tiers, and full language breadth is effectively gated to enterprise deals.
  • You have engineers who want maximum control and the cheapest possible orchestration layer, and you're comfortable assembling and pricing the rest of the stack yourselfVapi. It's model- and voice-agnostic by design, so you can chase the best model or cheapest voice without rebuilding the agent, and its API is widely cited as the cleanest of the category. Watch out for: the $0.05/min headline is only the orchestration fee — budget for $0.13–$0.33/min once a real model, voice and telephony are stacked on top.
  • Voice quality and latency on the actual call matter more to you than headline price, and you want built-in extras like a streaming knowledge base and pre-launch simulation testingRetell AI. Users consistently praise its natural-sounding calls and ~600ms turn-taking, and it ships an auto-syncing RAG knowledge base, batch outbound calling and an AI QA layer that simulates conversations before they go live. Watch out for: the same component-pricing pattern as Vapi — the $0.07/min headline realistically lands at $0.13–$0.20+/min once fully specced — plus some reviews flag a limited visual builder and slow support.
  • Voice quality and multilingual reach are the deciding factor, not the orchestration platform underneathElevenLabs Conversational AI. It starts from the industry's most-cited text-to-speech and instant-cloning models, natively speaks 70+ languages, and lets you reason with an ElevenLabs-hosted model for the lowest latency or bring your own LLM key, including open-source models. Watch out for: it covers the voice/LLM-orchestration layer only — you still assemble telephony, CRM/helpdesk workflow and observability yourself — and its free tier (15 minutes/month) is a demo, not a production trial.
  • You want to own the stack yourself — self-hosted, no vendor lock-in, full code visibility — and have the engineering time to assemble itLiveKit Agents. It's the only genuinely open-source (Apache 2.0) option here, proven at serious scale (it powers OpenAI's ChatGPT Advanced Voice Mode, plus xAI, Salesforce and Tesla), with a real free tier (1,000 agent-session minutes/month, no card). Watch out for: it's a framework, not a dashboard — you write and deploy code, and the free tier's small inference credit means real usage still means paying LiveKit or a model provider on top.
  • You want a proven, enterprise-scale vendor with nine years of production history and hundreds of named customers, and you're willing to start on a genuinely free but time-boxed self-serve tier before deciding on a larger engagementPolyAI. It's the oldest name in the category by years, with 200+ enterprise customers, thousands of live deployments across 75+ languages, and (since May 2026) a real self-serve on-ramp — free for two months from just an email address, via a no-code builder or a developer-focused ADK. Watch out for: what happens after the two-month free period is entirely undisclosed, and its largest deployments (like PG&E's 16M-call/year operation) still run through a separate, custom-quoted sales engagement.

LiveKit Agents: a fifth shape entirely, not a rival to the other four

Bland AI, Vapi, Retell AI and ElevenLabs Conversational AI are all closed-source SaaS — you rent an API or a dashboard, and the vendor's infrastructure is the product. LiveKit Agents doesn't compete on that axis at all: it's an Apache 2.0-licensed, open-source framework (11.5k+ GitHub stars) you write Python or TypeScript code against, and you can run the entire stack — speech-to-text, LLM orchestration, text-to-speech, telephony — on your own infrastructure, with no vendor able to change your pricing or shut off your access. LiveKit Cloud exists as an optional managed layer if you'd rather not run servers yourself, but nothing about the framework requires it. That openness is backed by real production weight: LiveKit is the company behind OpenAI's ChatGPT Advanced Voice Mode, and its other named customers include xAI, Salesforce and Tesla — this isn't an unproven side project, it's infrastructure already running at some of the largest voice-AI deployments anywhere. The honest tradeoff is the one every open-source framework carries versus a configured SaaS product: there's no hosted dashboard that gets a non-technical team a working agent in an afternoon, and the free tier's inference credit (~50 minutes of LiveKit-hosted model usage) is small enough that any real deployment needs a plan for ongoing LLM/STT/TTS costs on top of LiveKit's own agent-session-minute meter.

Vapi vs Retell AI: the closest real head-to-head

Of the four hosted platforms, Vapi and Retell AI are the actual rivals — both are developer-first, bring-your-own-model platforms with no self-hosted or fully-managed alternative, competing for the same technical-team budget. The difference is what's built in versus what you assemble yourself. Vapi's pitch is control and provider-agnosticism taken furthest: pick any STT, LLM and TTS, and coordinate multiple specialised agents ("squads") that hand off mid-call. Retell AI leans further into shipped tooling on top of the same bring-your-own-stack idea: an auto-syncing knowledge base, batch outbound calling and built-in simulation testing that Vapi doesn't list as native features, and user reviews credit it with the more natural-sounding voice output of the two. Neither publishes a true all-in price — both headline a base orchestration or pay-as-you-go rate that reviewers and the vendors' own docs confirm climbs several times higher once a real model, voice and telephony are added, so budgeting either one requires pricing the full stack, not the homepage number.

PolyAI: the pre-LLM enterprise incumbent, now with a self-serve on-ramp

PolyAI predates every other name on this list by years — founded in London in 2017, when Vapi, Retell AI, Bland AI, ElevenLabs Conversational AI and LiveKit Agents didn't yet exist. That history shows up as a scale gap: 200+ named enterprise customers, thousands of live deployments across 75+ languages in 25+ countries, and an $86M Series D (December 2025) that pushed its valuation to $750M, with NVIDIA's NVentures among the backers. Its Raven dialog model, trained on over a billion enterprise conversations, now also runs alongside third-party reasoning models (GPT-5, Claude, Gemini) since a May 2026 platform opening — the same event that ended PolyAI's sales-led-only era. For the first time, a team can sign up with just an email address and get two months of free access, either through a no-code Poly Agent Builder or a developer-focused Agent Development Kit (ADK) with self-serve API keys. That's a real change, not a repositioning — but it comes with an honest catch none of the other five share: PolyAI hasn't disclosed what its self-serve tier costs once the two-month free period ends, and its flagship deployments (PG&E's voice agents handle 16 million calls a year, with tens of thousands more during emergency spikes) still run through a separate, custom-quoted enterprise engagement rather than the new self-serve path. A vendor-commissioned but methodologically transparent Forrester study modelling 391% three-year ROI adds real, checkable numbers that most of the newer entrants here don't have a decade of customer relationships to produce.

One thing five of the six get right (and the one PolyAI complicates)

Vapi, Retell AI, Bland AI, ElevenLabs Conversational AI and LiveKit Agents all bill by usage — per-minute or a metered minute pool — not a flat per-seat licence, which fits a category built to replace or augment human phone work rather than seats at a desk. PolyAI's enterprise tier follows the same logic (a custom, sales-led per-minute rate bundling tuning, maintenance and support), but its new self-serve tier breaks the pattern entirely: it isn't metered at all, just free for a flat two-month window with no disclosed rate for afterward — the one pricing structure in this category that's neither usage-based nor a published subscription. None of the six publishes a genuine all-in per-minute number: Bland AI comes closest by bundling the LLM and voice into one rate, yet even its price varies by tier and platform fee; Vapi, Retell AI, ElevenLabs, LiveKit Agents and PolyAI's enterprise tier all headline a partial figure (an orchestration fee, a pay-as-you-go range, a subscription-plus-overage tier, a base agent-session-minute meter, or a custom-quoted bundle) that only becomes a real number once you've picked a model, voice and telephony provider and stacked their costs on top — or, for PolyAI's self-serve tier, once you find out what happens after month two. Any per-minute figure you see quoted for this category — on these listings or anywhere else — is a starting point for a spreadsheet, not a number to budget against directly.

FAQ

Which of these six owns its own infrastructure end to end?

Two do, in different ways: Bland AI runs its own model stack end to end as a hosted service and charges one all-inclusive rate, while LiveKit Agents is open-source (Apache 2.0) and fully self-hostable — you can run the entire stack on your own infrastructure, or use LiveKit's managed Cloud if you'd rather not. Vapi, Retell AI, ElevenLabs and PolyAI are all closed-source hosted platforms you can only rent — though PolyAI, like the other three, now lets you bring a third-party reasoning model (GPT-5, Claude or Gemini) rather than only its own proprietary Raven model.

Which is cheapest to start using?

By headline rate, Vapi's $0.05/min orchestration fee is the lowest number on the page. For an actually usable free tier, LiveKit Agents (1,000 agent-session minutes/month, no credit card) and PolyAI (fully free for two months, email signup only) are both far more generous than ElevenLabs' $0 tier (15 minutes/month, a demo allowance) — though PolyAI's free period is time-boxed with no published rate for afterward, unlike LiveKit's ongoing metered free tier. Headline price is not real cost for any of the six — Vapi, Retell AI, Bland AI, LiveKit and PolyAI's enterprise tier all bill separately for at least some of the model/voice/telephony components a full deployment needs, and ElevenLabs' free tier is a demo allowance, not a production budget.

What's the real difference between Vapi and Retell AI?

Mostly what ships built in versus what you assemble, not raw capability — see the head-to-head section above. Vapi pushes provider-agnosticism and multi-agent "squads" furthest; Retell AI adds a streaming knowledge base, batch calling and simulation testing on top of the same bring-your-own-stack model, and is more often praised for voice quality in reviews.

Is ElevenLabs Conversational AI a full voice-agent platform or just a TTS add-on?

It's a full agent layer — real-time transcription, an LLM turn and tool-calling, not just text-to-speech — but it covers the voice/orchestration layer only. Telephony beyond its native/Twilio connectors, CRM/helpdesk workflow and observability are left for you to assemble, the same limitation reviewers flag on the bring-your-own-stack platforms.

Is LiveKit Agents a turnkey product or something else entirely?

Something else entirely — it's an open-source framework and SDK, not a configured product. You write agent code (Python or TypeScript) that wires together a speech-to-text, LLM and text-to-speech pipeline, then deploy it yourself or to LiveKit Cloud. That's real engineering work the other platforms mostly hide behind a dashboard, but it also means no vendor can lock you in, raise prices unilaterally, or shut off access to your own infrastructure.

Which of these six is built for regulated industries?

Bland AI is built for this explicitly — it's the only hosted platform marketed at healthcare, finance and insurance with SOC 2, HIPAA and PCI DSS compliance plus self-hosted/on-premises deployment options built in from the start. PolyAI has the deepest track record actually running in regulated environments — named customers include a major US utility (PG&E) and a European bank (UniCredit), with SOC 2, HIPAA, GDPR and PCI DSS all shipping on the platform — though every regulated-scale deployment there still runs through PolyAI's sales-led enterprise path, not its new self-serve tier. LiveKit Agents is the third option worth evaluating: as a self-hosted open-source framework it gives a team full control over where data lives, though compliance certification and audit work then become the deploying team's own responsibility rather than a vendor-provided package.

How is PolyAI different from the other five voice agents?

PolyAI is the oldest name in the category by years (founded 2017, a Cambridge ML spinout) and the only one built around a large, named enterprise customer base (200+ customers) rather than a developer-first orchestration layer. Until May 2026 it was also the only one of the six with no self-serve access at all; it has since opened a free-for-two-months self-serve tier, though its largest deployments still run through a custom-quoted enterprise engagement, unlike the fully self-serve, always-on pricing Vapi, Retell AI, ElevenLabs and LiveKit Agents all publish.

Every agent above carries its own full researched listing — pricing tiers, sourced pros and cons, and a committed verdict on who it's for and who it isn't. Start with the voice agents category for the full set.