Summary
The Realtime API keeps a live session open between an application and OpenAI's gpt-realtime speech-to-speech models, which listen, reason, speak and call tools. The same endpoint family also runs dedicated live translation and streaming transcription sessions.
- Best for
- Engineering teams building a voice agent from scratch who want to go directly to OpenAI's own speech-to-speech model (no orchestration platform, no per-minute markup) and are comfortable being single-vendor on OpenAI.
- Not for
- Non-technical teams who want a working voice agent without writing code, or anyone who wants CRM integrations, call analytics or a dashboard out of the box: a managed platform like Vapi, Retell AI, Bland AI or Synthflow AI fits that need better.
- From
- $10/1m audio input tokensmetered on usage
- Free tier
- NoNew API accounts use prepaid billing with a $5 minimum credit purchase, and the rate-limit tables for gpt-realtime-2.1 and gpt-realtime-2.1-mini begin at Tier 1, whose qualification is $5 paid.
How much it does unattended
- Runs autonomously
- PartlyA session takes its own conversational turns and executes remote MCP tools itself, but your application must create the follow-up response after each tool call and owns function-tool execution.
The Realtime API doesn't create these follow-up responses automatically.
developers.openai.com/api/docs/guides/r… - Multi-agent
- YesA RealtimeAgent accepts handoffs like a text agent, so spoken workflows can branch across specialist agents in one session.
From there, attach tools, handoffs, and guardrails to the RealtimeAgent the same way you would attach them to a text agent.
developers.openai.com/api/docs/guides/v… - Agent permissions
- YesTool surface is constrained per session with allowed_tools and require_approval; an approval request becomes a conversation item your client must answer before the tool runs.
Keep the tool surface narrow with allowed_tools, and require approval for any action you would not auto-run.
developers.openai.com/api/docs/guides/r…
Which models it runs on
- Claude
- NoEvery realtime model the vendor enumerates is a first-party OpenAI gpt-realtime model; no Anthropic model appears in any published realtime model table.
gpt-realtime-2 is our state-of-the-art reasoning voice model for low-latency speech-to-speech applications.
developers.openai.com/api/docs/guides/r… - GPT
- YesThe endpoint runs OpenAI's own gpt-realtime speech-to-speech models with reasoning effort and function calling.
It supports speech-to-speech interactions with configurable reasoning effort, instruction following, and tool use for complex voice-agent workflows.
developers.openai.com/api/docs/models/g… - Open models
- NoThe realtime model choice is a closed set of OpenAI-hosted gpt-realtime models; no open-weight model is listed as supported on /v1/realtime.
gpt-realtime-2 is our state-of-the-art reasoning voice model for low-latency speech-to-speech applications.
developers.openai.com/api/docs/guides/r… - Model choice
- NoYou pick between the vendor's normal and mini realtime models; the vendor supplies the inference and no third-party model can serve a session.
The Realtime speech2speech models come in a “normal” size and a mini size, which is significantly cheaper.
developers.openai.com/api/docs/guides/r…
Where you use it
- In your editor
- NoThe documented surfaces are the three connection transports: WebRTC, WebSocket and SIP. No editor or IDE integration is published.
It supports audio and text inputs over WebRTC, WebSocket, or SIP connections and improves alphanumeric recognition over GPT-Realtime-2.
developers.openai.com/api/docs/models/g… - On the command line
- PartlyAn official openai command-line tool exists for the API, but the documented CLI workflows stop at request/response calls; realtime sessions are driven from your own WebSocket, WebRTC or SIP code.
Use the CLI for repeatable terminal workflows such as extracting structured data from files, generating images, creating speech, and composing API…
developers.openai.com/api/docs/libraries - In a browser
- YesWebRTC connects browser and mobile clients directly to a session, and the vendor runs a hosted Realtime Playground for testing prompts and measuring token usage.
For most browser voice agents, start with the Voice agents guide. It uses the Agents SDK with WebRTC for browser audio and can connect to server-side…
developers.openai.com/api/docs/guides/r…
Whose machine it runs on
- Self-hosted
- NoClients connect to OpenAI's hosted API server; every documented transport terminates on api.openai.com or sip.api.openai.com and no on-premise deployment is offered.
The Realtime API allows clients to connect directly to the API server via WebRTC or SIP.
developers.openai.com/api/docs/guides/r… - Open source
- PartlyThe service and the realtime models are proprietary; the client SDKs and the API's OpenAPI specification are published on GitHub.
You can also watch our OpenAPI specification repository on GitHub to get timely updates on when we make changes to our API.
developers.openai.com/api/docs/libraries
What it costs to run
- How it meters
- YesPay-as-you-go per token, with audio metered at 1 token per 100 ms of user audio and 1 token per 50 ms of assistant audio; translation and transcription sessions bill by audio duration instead.
Audio tokens in user messages are 1 token per 100 ms of audio, while audio tokens in assistant messages are 1 token per 50ms of audio.
developers.openai.com/api/docs/guides/r… - Free tier
- NoNew API accounts are prepaid with a $5 minimum credit purchase, and the published rate-limit tables for gpt-realtime-2.1 and gpt-realtime-2.1-mini start at Tier 1, which requires $5 paid.
Prepaid billing lets you purchase credits before using the API.
help.openai.com/en/articles/8264644-wha… - API access
- YesThe product is the API: your client opens a session on /v1/realtime and drives it with client events over WebRTC, WebSocket or SIP.
Voice-agent sessions use the standard Realtime API conversation lifecycle. The client connects to /v1/realtime, sends audio or text, and listens for…
developers.openai.com/api/docs/guides/r… - MCP server
- PartlyA Realtime session is an MCP client that calls remote MCP servers and OpenAI connectors itself, but the vendor publishes no MCP server that would let another agent query the Realtime API.
Use function tools when your application should execute the tool and return the result. Use MCP tools or built-in connectors when the Realtime API…
developers.openai.com/api/docs/guides/r… - Bring your own key
- NoRealtime sessions run only on OpenAI's own inference with OpenAI credentials; the AWS Bedrock path carries Responses API traffic and explicitly does not support WebSocket connections.
Amazon Bedrock provides Responses API-compatible inference for supported OpenAI models in supported AWS Regions.
developers.openai.com/api/docs/guides/a…
Buying it for a team
- A company can buy it
- YesBuying is organisation-first: an organisation holds billing and projects, and Admin APIs automate invites, users, projects and keys.
Organization: Your top-level account. Organization roles can grant access across all projects.
developers.openai.com/api/docs/guides/r… - Seat model
- NoBilling is per token, not per seat, and the vendor states rate limits are defined at organisation and project level rather than per user; no seat price, floor or ceiling is published.
Prices per 1M tokens unless noted.
developers.openai.com/api/docs/pricing - Pooled budget
- YesOne organisation balance with a monthly hard spend limit that returns 429 when reached, plus per-project spend alerts and project-level rate limits.
Use the Spend Limits endpoint to create or replace your organization's monthly hard spend limit.
developers.openai.com/api/docs/guides/a… - Admin controls
- YesAdmins constrain the agent per project: model allowlists or denylists, role-based permissions on /v1/realtime requests, IP allowlists and data-retention settings.
Use project model permissions to set an allowlist or denylist for a project.
developers.openai.com/api/docs/guides/a… - Audit log
- YesOrganisation audit logs are retrievable through the Admin API and its Audit logs endpoint.
Admin APIs let you automate organization management workflows such as user invitations, audit log review, project administration, API key management,…
developers.openai.com/api/docs/guides/a… - Single sign-on
- PartlySAML or OIDC single sign-on is configured for API Platform organisations through the tenant Admin Console, but the vendor tells you to confirm your API billing plan includes the SSO capability, and SCIM group sync is separate.
Check that your subscription or API billing plan includes the relevant SSO capability.
help.openai.com/en/articles/9534785
What happens to your code
- Opt out of training
- YesAPI data is excluded from model training by default; sharing is opt-in, and Zero Data Retention or Modified Abuse Monitoring can additionally remove content from the 30-day abuse-monitoring logs after approval.
Your data is your data. As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in…
developers.openai.com/api/docs/guides/y… - Data residency
- Partly/v1/realtime is listed for regional storage in all ten published regions and for regional processing in the United States and Europe, but residency is a sales-approved project option and residency endpoints carry a 10% price uplift for models released on or after March 5, 2026.
Data residency controls are a project configuration option that allow you to configure the location of infrastructure OpenAI uses to provide services.
developers.openai.com/api/docs/guides/y… - Getting out
- PartlyPurchased credits expire after a year and are non-refundable; on the data side /v1/realtime retains no application state and the Usage permission covers dashboard export.
Purchased credits expire after 1 year and are non-refundable.
help.openai.com/en/articles/8264644-wha… - Certifications
- YesThe API Platform is in scope for OpenAI's SOC 2 Type 2 report and its ISO/IEC 27001:2022 certificate, with 27017, 27018 and 27701 control sets and GDPR listed on the trust portal.
this report covered controls relevant to the Security, Availability, Confidentiality, and Privacy Trust Services Criteria for the API Platform,…
trust.openai.com
27 sourced claims on this page. Checked 2026-09-02. How we check.
Alternatives
Where we would send a reader instead, in that listing's own words.
Open-source Python and Node.js framework for voice AI agents, with LiveKit Cloud hosting, telephony and observability.
Hosted platform for building voice and chat agents, with telephony, SDKs and a choice of LLM.
Filed under
Sources and updates
The pages we read18
Every fact below comes from the vendor. Nothing here is independently corroborated yet.
- developers.openai.com/api/docs/guides/realtime-costs5 fields
- developers.openai.com/api/docs/guides/admin-apis4 fields
- developers.openai.com/api/docs/guides/realtime4 fields
- developers.openai.com/api/docs/guides/your-data4 fields
- developers.openai.com/api/docs/guides/realtime-mcp3 fields
- developers.openai.com/api/docs/guides/rbac2 fields
- developers.openai.com/api/docs/guides/realtime-models-prompting2 fields
- developers.openai.com/api/docs/libraries2 fields
- developers.openai.com/api/docs/models/gpt-realtime-2.12 fields
- developers.openai.com/api/docs/pricing2 fields
- help.openai.com/en/articles/8264644-what-is-prepaid-billing2 fields
- trust.openai.com/1 field
- developers.openai.com/api/docs/guides/amazon-bedrock1 field
- developers.openai.com/api/docs/guides/rate-limits1 field
- developers.openai.com/api/docs/guides/realtime-server-controls1 field
- developers.openai.com/api/docs/guides/voice-agents1 field
- developers.openai.com/api/docs/models/gpt-realtime-2.1-mini1 field
- help.openai.com/en/articles/95347851 field
Updates3
- Capabilities · Commercial terms · Free tier · Price
- Capabilities · Commercial terms · Free tier · Price
- One-liner
