# ElevenLabs Conversational AI alternatives

> Source: https://theagentsindex.com/elevenlabs-conversational-ai/alternatives (curated, quality-gated, re-verified)

ElevenLabs Conversational AI is a voice-quality-first agent layer where the buyer provides the LLM and inherits a strong default voice library. The seven alternatives below solve the same job at different points on the tradeoff between voice ownership, latency, transparency of the bill, and openness.

Some readers want a lower per-minute price. Others need self-hosting or an Apache-2.0 license. A few want one vendor running STT, LLM and TTS together rather than billing each layer separately. The ordering puts developer-platform parallels first and speech-infrastructure vendors' own agent layers at the bottom.

| # | Name | What it is | Why it's here | Pricing | API |
| --- | --- | --- | --- | --- | --- |
| 1 | [Vapi](https://theagentsindex.com/vapi.md) | The developer platform for voice AI agents, bring your own model, voice and telephony, and Vapi orchestrates the call. | Vapi is the developer platform overlap. BYO model and voice, $0.05/min orchestration fee on top of whatever model, voice and telephony the team picks. ElevenLabs positions itself around its own voice library and instant cloning. Vapi fits when voice identity matters less than the wire-level call plumbing. | $0.05 / min | Yes |
| 2 | [Retell AI](https://theagentsindex.com/retell-ai.md) | Build phone-call voice agents on Retell's turn-taking layer, plug in your own LLM, voice, and telephony, ship over voice, chat, or SMS. | Also a BYO-model developer platform, with a turn-taking layer as the main differentiator. Pricing starts at $0.07/min for AI voice agents, then accumulates a Retell voice-infrastructure fee, a TTS fee, a telephony fee and the LLM cost on top. Anyone choosing Retell over ElevenLabs is usually optimizing for response cadence, not voice cloning. | $0.07–$0.31 / min | Yes |
| 3 | [Bland AI](https://theagentsindex.com/bland-ai.md) | Enterprise voice AI that owns its own model stack, priced all-in per minute with no model pass-throughs. | Bland prices all-in at $0.14/min with no model pass-throughs, which makes the bill easier to forecast than ElevenLabs' credit-pool-plus-overage structure. The trade is voice catalog. Bland runs its own model stack and a smaller voice library, so the move makes sense for teams that care more about predictable unit economics than about ElevenLabs' 10,000-voice range. | $0.14/min / minute | Yes |
| 4 | [Cartesia](https://theagentsindex.com/cartesia.md) | The Stanford SSM-research spinout behind Sonic (TTS) and Ink (STT) — Line, its open-source voice-agent SDK, connects to 100+ LLMs and starts at a genuine $0/month, but the agent layer itself isn't self-hostable. | Cartesia ships Line, an open-source voice-agent SDK, on top of its Sonic TTS and Ink STT. Agent minutes meter at $0.06/min regardless of tier, with a model-credit pool bought separately. The Stanford SSM pedigree and the open-source SDK make Cartesia the choice for latency-sensitive buyers who do not need ElevenLabs' voice cloning. | $0 / month | Yes |
| 5 | [Deepgram Voice Agent API](https://theagentsindex.com/deepgram-voice-agent-api.md) | The speech-infrastructure vendor's own voice-agent API. Managed STT, LLM and TTS over a single WebSocket, $4.50/hr flat, with 6 LLM providers and self-hosted or VPC deployment on the same product. | Deepgram bundles STT, a configurable LLM and Aura-2 TTS over one WebSocket at a flat $4.50/hr. ElevenLabs' 10,000-voice catalog is replaced here by a much smaller Aura-2 library and the option to fully replace every layer with a bring-your-own component. The right alternative for buyers who want one vendor running the whole turn pipeline. | $4.50/hr Standard / Standard to Custom (BYO LLM and TTS) | Yes |
| 6 | [LiveKit Agents](https://theagentsindex.com/livekit-agents.md) | The Apache 2.0 voice-agent framework. 13.1k GitHub stars, self-hostable, and the only open-source pick in this category with fully published pricing. | The only Apache-2.0 entry on this list. The Build tier is $0/month with 1,000 agent-session minutes included, and the framework runs self-hosted with no platform fee. LiveKit is the alternative for engineering teams that refuse a managed platform, and it is also the place where voice is one of many real-time media types the same SDK handles. | $0 / month | Yes |
| 7 | [OpenAI Realtime API](https://theagentsindex.com/openai-realtime-api.md) | OpenAI's own speech-to-speech model API for voice agents: WebRTC, WebSocket or native SIP telephony, billed purely on audio tokens with no platform fee. | OpenAI's own speech-to-speech model, billed purely on audio tokens and usable over WebRTC, WebSocket or SIP telephony. The flagship gpt-realtime model sits at roughly $0.06 to $0.11/min depending on cache hit rate, with the smaller mini model at roughly a third of that. The right pick for teams that want OpenAI as the LLM and would rather not layer a separate agent platform on top. | $32 / $64 audio / per 1M tokens | Yes |

## Methodology

Every entry here is a Published voice-agent platform that can be deployed as the only system on the call. It accepts an inbound voice conversation, talks back, and either resolves the call or hands it off. Excluded platforms where voice is one channel among many, tools where the agent layer is not separately licensable, and enterprise voice-AI products sold only under multi-year custom contracts.

Ranked by closeness of substitute to ElevenLabs Conversational AI. Developer-platform parallels come first, then specialist voice vendors with their own model stacks, then the speech-infrastructure and LLM vendors' own agent layers. Prices are the entry tier as published on each vendor's own page, re-verified weekly.
