Realtime API¶
Hold a live, bidirectional speech-to-speech conversation over a single WebSocket, through the OpenAI Realtime API shape. Audio flows in both directions on the same connection: send the caller's speech as it is captured, and receive the model's spoken answer as it is generated — no request/response round trip per turn.
At a glance¶
- The same client and server event vocabulary as the OpenAI Realtime API —
client.realtime.connect(model=...)works by changing the base URL — see Feature compatibility. - Ephemeral client secrets valid 10 to 7,200 seconds — mint an
ek_...credential server-side and hand it to an untrusted browser or mobile client; your API key never leaves your backend — see Ephemeral client secrets. - A signed token, not a server-side record — any instance behind a load balancer verifies a secret minted by any other, with no shared session store — see Ephemeral client secrets.
- 24 kHz PCM by default, or G.711 at 8 kHz —
audio/pcmuandaudio/pcmainteroperate directly with telephony and SIP media — see Feature compatibility. - WebSocket always, WebRTC by operator opt-in, SIP never terminated here —
POST /v1/realtime/callsanswers404untilREALTIME_WEBRTC_ENABLEDis set — see Transports. - The model calls your own functions — declare them on the session, run the call, hand back the result, and the answer is spoken with it — see Function calling.
- One session lasts at most 8 minutes — the connection then closes with reason
session_expired, and reconnecting continues the conversation — see Session lifecycle and limits.
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI(api_key="YOUR_API_KEY", base_url="https://your-gateway/v1")
async def main() -> None:
async with client.realtime.connect(model="amazon.nova-2-sonic-v1:0") as connection:
await connection.conversation.item.create(
item={
"type": "message",
"role": "user",
"content": [
{"type": "input_text", "text": "Say hello in one sentence."}
],
}
)
await connection.response.create()
async for event in connection:
if event.type == "response.output_audio_transcript.delta":
print(event.delta, end="", flush=True)
elif event.type == "response.done":
break
asyncio.run(main())
Endpoints¶
| Endpoint | Method | What It Does | Powered By | MCP Tool |
|---|---|---|---|---|
/v1/realtime/client_secrets | POST | Mint a short-lived, signed client secret carrying a session configuration | Amazon Bedrock | openai_realtime_client_secret |
/v1/realtime?model=<id> | WS | Open a live, bidirectional speech-to-speech session | Amazon Bedrock | Not applicable — a persistent connection |
/v1/realtime/calls | POST | Trade a WebRTC SDP offer for an SDP answer, opening a call the gateway terminates — opt-in | Amazon Bedrock | Not applicable — negotiates a media connection |
/v1/realtime/calls/{call_id}/hangup | POST | End an active WebRTC call | Amazon Bedrock | Not applicable — controls a media connection |
WebSocket always — WebRTC opt-in, SIP never
Every deployment serves the WebSocket. POST /v1/realtime/calls answers WebRTC offers only when the operator enables REALTIME_WEBRTC_ENABLED — it needs a UDP media path the default deployment does not have — and answers 404 otherwise. Inbound SIP is never terminated here: the SIP-only verbs (accept, reject, refer) answer a clean 400 naming what serves telephony instead.
Transports covers the whole picture: how a browser connects, what enabling WebRTC entails, and how to put a media framework or a phone line in front of this deployment.
Feature compatibility¶
| Feature | Status | Notes |
|---|---|---|
| Client Events | ||
session.update | Voice, instructions and audio formats are fixed once the conversation opens — see below | |
input_audio_buffer.append | Base64-encoded audio in the session's configured input format; at most 4 MiB per event, so send it in chunks as it is captured | |
input_audio_buffer.commit | Required to end a turn when turn_detection is null; at most 5.7 MB of audio may wait for one | |
input_audio_buffer.clear | Discards buffered, not-yet-committed audio | |
conversation.item.create | Text items and function_call_output items, whether they add to the conversation or replay its history. An audio item is refused with a clear error — send speech through input_audio_buffer.append — and so is a function_call item; see Function calling | |
conversation.item.truncate | Answered with conversation.item.truncated — see below | |
conversation.item.retrieve | Answered with conversation.item.retrieved, carrying the item's role, status and transcript; audio is not retained, so the item carries no audio field | |
conversation.item.delete | Answered with conversation.item.deleted; the item stops being addressable, and the model keeps its own memory of the conversation | |
response.create | Ends any open turn and starts the model answering; its per-response response payload is ignored, and the session's own configuration serves every answer | |
response.cancel | Ends the answer in progress with status: "cancelled"; what the model keeps speaking is dropped rather than reported | |
output_audio_buffer.clear | Acknowledged with output_audio_buffer.cleared | |
| Server Events | ||
session.created / session.updated | Sent on connect and after every accepted session.update | |
input_audio_buffer.speech_started / .speech_stopped | Server-side voice activity detection only | |
input_audio_buffer.committed / .cleared | ||
conversation.item.added / .done | Sent for every item: a written one, the caller's committed audio, and each answer — the answer's .done precedes its response.done | |
conversation.item.created | Sent beside conversation.item.added for a written item and for each function call the model makes, for clients predating the added/done pair. Not sent for the caller's own committed turn or for the answer's own item, which are announced with the added/done pair alone | |
conversation.item.truncated / .retrieved / .deleted | Answers to the matching client event | |
conversation.item.input_audio_transcription.delta / .completed | Only when audio.input.transcription is set on the session | |
conversation.item.input_audio_transcription.failed | Sent instead of a transcript when a caller turn could not be read, so it is not mistaken for a caller who said nothing; only when audio.input.transcription is set | |
conversation.item.input_audio_transcription.segment | Not emitted — a transcript carries no per-speaker segments or timings | |
response.created / response.done | Both carry the whole response object — see below; response.done adds the answer's token usage | |
response.output_item.added / .done | ||
response.content_part.added / .done | ||
response.output_audio.delta / .done | Spoken answers only | |
response.output_audio_transcript.delta / .done | Spoken answers only | |
response.output_text.delta / .done | Text-only answers (output_modalities: ["text"]) | |
response.function_call_arguments.delta / .done | Sent for every function call; the arguments arrive whole, in one delta | |
output_audio_buffer.cleared | ||
error | Non-fatal for a rejected event; terminal (closes the socket) for a fatal one | |
rate_limits.updated | Not emitted | |
input_audio_buffer.timeout_triggered | Not emitted — it reports an idle_timeout_ms that is not available | |
output_audio_buffer.started / .stopped | Not emitted, on the WebRTC transport either | |
| Audio Formats | ||
audio/pcm | Default — 24 kHz, 16-bit, mono, little-endian | |
audio/pcmu, audio/pcma | G.711 at 8 kHz, for telephony interoperability | |
| Independent input/output formats | Configured separately under audio.input.format / audio.output.format | |
| Turn Detection | ||
| Server-side voice activity detection | Default; ends each turn automatically | |
Manual turns (turn_detection: null) | End each turn yourself with input_audio_buffer.commit | |
turn_detection.type: "semantic_vad" | Accepted and served as server_vad — a turn ends on silence, not on what was said | |
threshold, prefix_padding_ms, silence_duration_ms, idle_timeout_ms, eagerness | Accepted and ignored — detection sensitivity is not tunable | |
create_response, interrupt_response | Accepted and ignored — a detected turn always starts a response, and interruption is the model's own decision | |
| Barge-in (caller speaks over the answer) | Handled by the model itself | |
| Voices | ||
| OpenAI voice names | alloy, ash, ballad, cedar, coral, echo, marin, sage, shimmer, verse — each served by the model's own nearest voice, so the timbre is not the upstream one | |
| Any other voice name | Passed through to the model as given, so a model voice can be named directly | |
| Custom voice object | audio.output.voice also accepts {"id": "…"}, on session.update and on the session a client secret carries; the id names the voice, as the plain string does | |
| Tools | ||
tools (type: "function") | Declared before the conversation opens and fixed for the rest of it — see Function calling | |
tools (type: "function") without a name | Accepted and ignored — nothing can be called by no name, so the entry is dropped and the rest of the session stands | |
tools (type: "mcp") | Refused — a session calls the functions its client runs, never a remote MCP server. POST /v1/realtime/client_secrets refuses it with 400, and a session.update with an error | |
tool_choice | auto, none, required and a named function; required and a named function make every answer start with a call | |
tool_choice (type: "mcp") | Accepted and ignored — no remote MCP server is ever attached, so the session behaves as auto | |
parallel_tool_calls | Accepted and ignored — how many tools one answer calls is the model's own decision | |
| Not Available | ||
POST /v1/realtime/calls (WebRTC) | Opt-in: served when REALTIME_WEBRTC_ENABLED is set and the deployment has a UDP media path; answers 404 otherwise | |
SIP (accept, reject, refer call verbs) | Inbound SIP is not terminated by the gateway, permanently — the verbs answer 400; see Transports for the telephony route | |
prompt (prompt templates) | Accepted and ignored | |
reasoning, tracing, truncation | Accepted and ignored | |
include | Accepted and ignored — no extra output fields are available | |
audio.input.noise_reduction | Accepted and ignored — incoming audio is not filtered | |
audio.input.transcription.model / .language / .prompt | Accepted and ignored — the transcript comes from the session's own model, which detects the language and takes no vocabulary hint. Setting the transcription object at all is what turns the events on | |
audio.output.speed | Accepted and ignored — the spoken answer is not time-scaled |
Legend:
- Supported — Fully compatible with OpenAI API
- Conditional — Depends on session configuration
- Partial — Supported with limitations
- Unsupported — Not available in this implementation
- Extra Feature — Enhanced capability beyond OpenAI API
Voice, instructions and audio formats are fixed once the conversation opens¶
The model's voice, its system instructions, both audio formats and the session's tools are set when the conversation with the model opens, and cannot change for the rest of that session. The conversation opens on the first thing sent into it — the first input_audio_buffer.append under the default server voice activity detection, or the first input_audio_buffer.commit, response.create or conversation.item.create under manual turns — which is well before the model has answered anything.
Send session.update with these settings before sending anything else, or open a new session to change them. Afterwards, a session.update touching only other fields (turn_detection, max_output_tokens, transcription settings, and so on) is still accepted; one that would change voice, instructions, an audio format, tools or tool_choice is refused with an error event.
Answering a written turn¶
conversation.item.create carrying an input_text part adds the text to the conversation without starting an answer. Follow it with response.create, and the model answers it exactly as it answers a spoken one — the same response.output_audio.delta chunks and the same transcript — so a written nudge into a voice session ("the caller has been on hold", "wrap up now") needs no second channel.
What a response object reports¶
response.created and response.done carry the same response object, and every field the upstream API sends is present on both — a voice framework validates each frame against its own models, and a missing field is the event never arriving rather than a cosmetic difference.
| Field | What it carries |
|---|---|
status_details | null while the answer is in progress and once it has completed. An answer that was stopped reports status: "cancelled" with the reason it stopped: {"type": "cancelled", "reason": "turn_detected"} when the caller spoke over it, {"type": "cancelled", "reason": "client_cancelled"} when response.cancel ended it. The conversation item of a stopped answer settles as incomplete, since what was said before the stop still stands |
conversation_id | The conversation the answer was added to — one per session, so every answer of a session names the same one |
output_modalities | ["audio"], or ["text"] when the session asked for text-only answers |
max_output_tokens | The session's max_output_tokens, or "inf" when it sets none |
audio | The session's effective output format and voice; voice is null when the session named none and the model answered in its own |
metadata | Always null — an answer carries no metadata, since response.create takes no per-response configuration |
output, usage | The answer's conversation item, and the tokens it used (on response.done) |
Truncating an answer the caller spoke over¶
The model generates speech faster than it is played, so a caller who interrupts has heard less of the answer than was sent. Send conversation.item.truncate with the item's id, content_index: 0 and the audio_end_ms your player actually reached; the session cuts its record of that item to what was heard and answers conversation.item.truncated.
- The item's transcript is removed whole, not trimmed: nothing aligns a transcript to a position in the audio, and leaving text the caller never heard in the record is the failure this event exists to prevent.
audio_end_mspast the end of the item's audio, an item that is not an assistant message, and an item this session never sent are each refused with anerror.- What the model itself remembers of the answer is the model's own; truncation aligns the record this session reports through
conversation.item.retrieved.
Function calling¶
Declare functions on the session, and the model asks for one whenever it needs what only your application knows — an account balance, a booking, the state of a device. The call arrives as a finished answer, so the client can run it immediately; the result goes back as a conversation item, and the model speaks its reply with what the function returned.
await connection.session.update(
session={
"type": "realtime",
"tools": [
{
"type": "function",
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {"location": {"type": "string"}},
"required": ["location"],
},
}
],
"tool_choice": "auto",
}
)
A call is reported as its own conversation item and its own response:
response.output_item.addedcarries afunction_callitem with the function'snameand thecall_idto answer it by.conversation.item.createdandconversation.item.addedannounce the same item.response.function_call_arguments.deltathen.donecarry the arguments as a JSON string. They arrive whole, in one delta.response.doneends that answer withstatus: "completed", thefunction_callitem in itsoutput.- Send the result back with a
function_call_outputitem naming the samecall_id:
await connection.conversation.item.create(
item={
"type": "function_call_output",
"call_id": call_id,
"output": '{"temperature_c": 14, "condition": "rain"}',
}
)
The model starts answering as soon as the function_call_output arrives — step 5 upstream, an explicit response.create, is not needed here. Sending one anyway is safe: it ends the open turn exactly as it would have, and the answer is still a single response.
Things to know before wiring an agent to it:
- Answer every call. A model waiting for a result says nothing else for the rest of the session, so return one even when the function failed — an
{"error": "..."}payload is a usable answer, silence is not. outputis free text, and a JSON object travels best. Anything that is not one is carried as theresultfield of one, so a function returning structured data should return it as JSON.- Declare the tools before the conversation opens — with the voice, the instructions and the audio formats (above). A
session.updatechangingtoolsortool_choiceafterwards is refused with anerror. - One call per answer. Each call ends the answer that produced it, and anything the model says afterwards is a new response; a client tracking responses sees more of them in a session that calls functions.
- Any
call_idis accepted. Afunction_call_outputis carried to the model whatever it names, so a client may replay a conversation's history into a fresh session, or resend an answer after reconnecting past the session cap, and the model itself decides what to do with it. Nothing is refused for naming a call this particular connection did not produce.
A cancelled answer's calls are answered for you
A model left waiting for a result never speaks again, so a call that arrives while its answer is being suppressed — after response.cancel, before the model has stopped — is answered on your behalf with {"error": "The answer was cancelled."}, and neither the call nor that answer is reported to you. Without it the session would sit silent until its own cap; with it, the model has seen one turn your application never authorised, which can surface in what it says next. Cancel between turns rather than mid-call where the flow allows it.
A function_call item cannot be written by the client
Upstream's creatable item union includes function_call, and this API refuses it with an error naming that reason: the conversation holds the calls the model itself made, and one written in from outside cannot be presented to the model as its own. Replay the results with function_call_output items, which are accepted whatever call_id they name, and put anything else the model needs to know into a text item or the session instructions.
The functions run in your application
A tool is a name, a description and a JSON Schema — the deployment never runs it, never reaches the network for it, and never sees more of it than the result you hand back. Remote MCP servers (tools entries of type: "mcp") are refused for the same reason: nothing here calls out to a third party on the caller's behalf. POST /v1/realtime/client_secrets refuses them too, rather than minting a secret carrying a session no connection could open.
Guardrail coverage¶
When the deployment configures an Amazon Bedrock guardrail, a realtime session is checked per turn: what the caller said (as the model transcribes it) as INPUT, and each completed answer as OUTPUT. A blocked turn ends the session with a terminal error event and close code 3000.
Unlike a request/response route, the check cannot come before the content reaches the client: the model's speech is streamed while it is being generated and its transcript is only complete once the answer is over, so a blocked answer may already have been partly heard when the session ends. Written items sent with conversation.item.create are checked as INPUT before they reach the model, as on every other route. So is what a function returns; the arguments the model asks a function to run with are not checked, since they are a request for data rather than content spoken to the caller.
Models¶
Every deployment's catalog differs, and a model's own name is never guaranteed stable across accounts. Find which models serve this route:
curl "$BASE/search_models?route=openai_realtime" \
-H "Authorization: Bearer $OPENAI_API_KEY"
Pass the returned model ID as model on the WebSocket URL, or in the session.model field of an ephemeral secret's configuration. See the Search Models API for the full filter syntax.
session.model does not accept a wildcard pattern; the WebSocket's model does
POST /v1/realtime/client_secrets fixes the model into the signed token before a connection exists, so session.model must name an exact model — a wildcard pattern is rejected. The model query parameter of WS /v1/realtime itself has no such constraint and accepts a pattern.
Authentication¶
Open the WebSocket with a credential, carried in whichever of three channels the client can use:
| Client | Credential carrier |
|---|---|
| Server-side SDKs | Authorization: Bearer <api key or ephemeral secret> header |
| Other gateway clients | x-api-key: <api key or ephemeral secret> header |
| Browser (cannot set custom WebSocket headers) | Sec-WebSocket-Protocol list entry openai-insecure-api-key.<credential> |
Any credential the deployment accepts on its HTTP routes works here — its own API key, a tenant API key where tenants are configured, an Amazon Cognito user pool access token where a pool is — and so does an ephemeral client secret (ek_...). A handshake made with a tenant key counts as one request against that tenant's rate limits, and the tenant's endpoint and model restrictions apply to /v1/realtime as to any other route. The model query parameter (/v1/realtime?model=<model id>) selects the model serving the session; it may be omitted when the credential is an ephemeral secret whose session configuration already names one. See Authentication & Security for how the API key itself is configured.
A refused credential is not an HTTP status
The WebSocket upgrade always completes first, so a rejected or expired credential is not answered with 401/403. The connection opens, the first and only event is a terminal error with code: "invalid_api_key", and the socket is then closed with close code 3000 and reason invalid_request_error.invalid_api_key — the same shape the upstream API uses. Instrument the error event and the close code, not the handshake status.
Ephemeral client secrets¶
POST /v1/realtime/client_secrets mints a short-lived credential — a value starting with ek_ — that carries a session configuration. Hand it to a browser or mobile client so it can open a session directly, without ever holding the deployment's own API key.
curl -X POST "$BASE/v1/realtime/client_secrets" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"expires_after": {"anchor": "created_at", "seconds": 60},
"session": {
"type": "realtime",
"model": "amazon.nova-2-sonic-v1:0",
"instructions": "You are a helpful support agent."
}
}'
{
"value": "ek_...",
"expires_at": 1731000060,
"session": {
"type": "realtime",
"model": "amazon.nova-2-sonic-v1:0",
"instructions": "You are a helpful support agent.",
"audio": { "...": "..." }
}
}
expires_after.secondsaccepts 10 to 7,200 seconds, defaulting to 600 (10 minutes) when omitted. This bounds how long the secret can be used to open a session — a session already opened with it keeps running for its own session limit.sessionaccepts the same configuration a client would otherwise send in asession.updateevent; it is applied to every session opened with the secret. It also accepts{"type": "transcription"}, which opens sessions that only transcribe the caller and never answer out loud — the only way to ask for one, since the socket itself takes no session type. Such a session still producesresponse.*text events, and still costs what a spoken session costs.- By default the carried configuration is a default, not a constraint: the client may name another model on the
?model=query string and change the configuration with its ownsession.update, as it can upstream. SetREALTIME_ALLOW_SESSION_OVERRIDE=falseto make the model, theinstructionsandmax_output_tokensthe secret was minted with final — a mismatching?model=is then refused at connect, and asession.updatechanging one of them answers anerror.
What a secret grants until it expires
A secret cannot be revoked: rotating REALTIME_CLIENT_SECRET_KEY invalidates every outstanding one at once, and nothing else does. Until then it may open any number of concurrent sessions, each billed to the deployment — so keep expires_after.seconds as short as the flow allows.
Its payload is signed, not encrypted: whoever holds the secret can read the session configuration it carries. Nothing confidential belongs in instructions.
Stateless, and signed
Nothing is stored server-side: the secret is the session configuration plus a signature, so any instance behind a load balancer verifies a secret minted by any other — no shared session store, no sticky routing required.
The signing key is derived from the deployment's configured API key by default. When the deployment runs with no API key configured at all, the signing key falls back to a random value generated per process: minted secrets then only verify on the instance that minted them, and stop working once a request reaches a different one. Set realtime_client_secret_key explicitly to fix a key shared by every instance regardless of the API key configuration.
Transports¶
Upstream offers a realtime session over three transports — WebSocket, WebRTC and SIP. This API always serves the WebSocket; WebRTC is an operator opt-in; SIP is never terminated here. POST /v1/realtime/calls trades an SDP offer for an answer once REALTIME_WEBRTC_ENABLED is set — see WebRTC calls terminated by the gateway for what that transport delivers and what it demands of the deployment — and answers 404 otherwise.
A browser connects to that same WebSocket¶
There is no separate browser transport to be missing. The page opens wss://<host>/v1/realtime?model=<id> itself, carrying an ephemeral client secret in the way Authentication describes, so the deployment's own API key never leaves your backend. It is two steps — mint the secret server-side, connect with it client-side — and the browser example below is both of them.
What the page owns in exchange is the media. Capturing the microphone, resampling it to the session's input format and playing back the response.output_audio.delta chunks are its own work, because a WebSocket carries the audio bytes handed to it and nothing else: no jitter buffer, no packet-loss concealment, no echo cancellation. On a good network that is unremarkable; on a lossy one it is audible, and it is the reason to reach for a media stack rather than write one.
WebRTC calls terminated by the gateway¶
With REALTIME_WEBRTC_ENABLED set — and the webrtc optional dependencies installed, which the container images ship — the gateway terminates the whole WebRTC media path itself: ICE, DTLS-SRTP and Opus. A browser (or any WebRTC client) posts its SDP offer and connects directly, with nothing in between:
POST /v1/realtime/callsaccepts the offer as a rawapplication/sdpbody — with?model=<id>on the query string, exactly as upstream's browser flow — or asmultipart/form-datawith ansdpfield and an optionalsessionJSON field. Either encoding authenticates with an ephemeral client secret or the deployment's credentials; a secret whose session is locked refuses asessionfield. A JSON body is refused withunsupported_content_type, which is what the upstream endpoint answers too. The response is201with the SDP answer as its body and the call's identifier in theLocationheader (/v1/realtime/calls/rtc_...).- Audio rides the media tracks as Opus, both directions — no base64, no
response.output_audio.deltaevents. Events ride a data channel the client opens under the labeloai-events, in exactly the WebSocket vocabulary:session.createdarrives on it once the channel opens,session.update,response.createand the rest work unchanged. The session's audio formats are fixed by the media negotiation, so asession.updatechanging them is refused with anerror. - Playback is the gateway's, so barge-in is too. The model generates speech faster than it is played, and on a call the unplayed tail sits in the gateway rather than in your player. A caller who starts speaking over it —
input_audio_buffer.speech_started— stops hearing it immediately, and the drop is reported asoutput_audio_buffer.cleared, as it is upstream. An answer ended byresponse.cancelor interrupted mid-generation stops the same way. POST /v1/realtime/calls/{call_id}/hangupends the call; the SDP answer'sLocationheader is where the identifier comes from.- A sideband WebSocket —
WS /v1/realtime?call_id=<id>, API credentials only, never an ephemeral secret — observes the call's server events and may send client events into its session, mirroring upstream's monitoring channel. - Call control follows the credential that opened the call. A call opened under a tenant API key can be ended or observed only with that tenant's key or the deployment's own credentials; any other caller is answered the same
404as an unknown call.
What a WebRTC call demands of the deployment
- A UDP path to the exact instance that answered. ICE negotiates ephemeral UDP ports directly to the server process; an HTTP(S) load balancer cannot carry them. The Terraform module's WebRTC media mode provisions the public task IP and UDP ingress this needs, off by default. Behind 1:1 NAT, set
REALTIME_WEBRTC_STUN_SERVERso the gateway advertises its public address; for callers on UDP-blocking networks, run a TURN relay (for example coturn) and setREALTIME_WEBRTC_TURN_SERVER,REALTIME_WEBRTC_TURN_USERNAMEandREALTIME_WEBRTC_TURN_PASSWORD— the three are required together, and a deployment that sets only some of them refuses to start, naming them. AWS offers no managed TURN. - The caller must offer a publicly routable candidate. An SDP offer names the addresses the gateway sends its ICE checks to, so candidates on addresses that are not globally routable — private, loopback, link-local — are dropped, and an offer left with none is refused with
invalid_offer. Hostname and mDNS (.local) candidates are always dropped. For callers that legitimately share the deployment's network, setREALTIME_WEBRTC_ALLOW_PRIVATE_CANDIDATES. - One instance, or routed call control. A call lives in the memory of the instance that answered its offer.
hangupand the sideband WebSocket answer404on any other instance; run a single instance, or route call-control requests to the answering instance yourself. - The 8-minute session cap applies to calls too. Amazon Nova Sonic ends a session at 480 seconds, so a call hard-stops at 8 minutes with the connection torn down — a real limit for the phone-length conversations WebRTC invites.
- Scale-in, deployments and Spot interruption end live calls. The media path cannot drain: an instance that stops mid-call drops it.
Put WebRTC or a phone line in front of the gateway¶
The route to a browser peer connection or a phone call does not have to run through POST /v1/realtime/calls, and for SIP it never does. The two frameworks most voice agents are already built on — LiveKit Agents and Pipecat — terminate WebRTC themselves (and SIP, through their telephony transports), and reach the model over exactly the WebSocket this API serves. The caller speaks WebRTC or SIP to the framework; the framework speaks this API. Pointing one at this deployment asks no more of it than the rest of this gateway does: the base URL, the deployment's API key, and the model name only where the name differs.
LiveKit Agents takes an HTTP base URL and derives the WebSocket from it, exactly as the official SDK does:
from livekit.agents import AgentSession
from livekit.plugins import openai
session = AgentSession(
llm=openai.realtime.RealtimeModel(
model="amazon.nova-2-sonic-v1:0",
base_url="https://your-deployment.example.com/v1",
api_key="YOUR_API_KEY",
)
)
base_url also reads from the OPENAI_BASE_URL environment variable, and api_key from OPENAI_API_KEY. A base URL ending in /v1 has /realtime appended for you; a deployment served under a non-default OPENAI_ROUTES_PREFIX has to name the full path itself.
Pipecat takes the WebSocket URL whole, /v1/realtime included:
import os
from pipecat.services.openai.realtime.llm import OpenAIRealtimeLLMService
llm = OpenAIRealtimeLLMService(
base_url="wss://your-deployment.example.com/v1/realtime",
api_key=os.environ["OPENAI_API_KEY"],
settings=OpenAIRealtimeLLMService.Settings(model="amazon.nova-2-sonic-v1:0"),
)
On a telephony leg, set the session's audio formats to audio/pcmu or audio/pcma — G.711 at 8 kHz is what a phone call already carries, so nothing resamples it twice.
Check the constructor against the version you install
The parameters above are as shipped in livekit-plugins-openai 1.6.10 and pipecat-ai 1.7.0. Both projects have renamed this integration before — Pipecat's service moved out of pipecat.services.openai_realtime_beta, and LiveKit does not list base_url on its own parameters page — so read LiveKit's OpenAI realtime plugin and Pipecat's OpenAI realtime service for the release you pin.
Why WebRTC is a different kind of endpoint¶
An HTTP request that returns an SDP answer is the small, visible part of WebRTC. The rest is a media transport in its own right: a UDP path negotiated separately from the HTTPS connection, carrying encrypted RTP, with its own address discovery, NAT traversal and congestion control. Serving it means terminating that media path, not routing one more request — and the two are not the same thing to build, nor the same thing to put an ingress in front of. SIP has the same shape under different names: signalling on one connection, audio on another.
That is also why a framework remains the shorter path for anything the gateway-terminated transport does not cover — telephony, more than one instance, calls past 8 minutes, hostile networks without your own TURN. LiveKit and Pipecat already own a media stack, already run wherever your users are, and already speak this API on the other side. WebRTC and SIP need their own ingress covers what terminating media costs a deployment, whichever process does it.
Billing¶
A realtime session bills audio and text tokens continuously, in both directions, for as long as the connection is open — not per request. Usage is reported per answer, in each response.done event, and recorded in the gateway's usage log the same way, so a session that drops mid-conversation still accounts for everything spoken before it. Speech tokens are priced well above text tokens by AWS, and the two are recorded and priced separately here. See Cost Management for how usage becomes cost.
A silent session costs what a spoken one costs
The backend has no mode that answers without speaking: the model synthesises every answer, and asking for output_modalities: ["text"] or a transcription session suppresses delivery of that audio, not its generation. The speech tokens are produced, counted in the usage event and billed. Choose either for the event stream it gives a client, never to lower the bill.
Limits and behaviour to know¶
What a session refuses and how it ends. The event-level gaps are in the feature compatibility table, and what a WebRTC call additionally demands is under WebRTC calls.
Session lifecycle and limits¶
- Duration cap — a session lasts at most 8 minutes. When it is reached, the server closes the connection with WebSocket close code
1000and reasonsession_expired; reconnect to continue the conversation. A WebRTC call is under the same cap: at 8 minutes its session ends and the peer connection is torn down. - Conversation ended by the model side — when the model ends the conversation itself, the connection closes with close code
1000and reasonsession_ended. Normal, and reconnecting starts a new session. - Server shutdown — a session still open when the deployment shuts down is closed with close code
1001and reasonserver_shutdown. - Fatal errors — a fatal error sends a terminal
errorevent, then closes the connection with close code3000, whose reason is<error type>.<error code>(e.g.invalid_request_error.model_not_found). - Event size — a single client event may carry at most 4 MiB, base64 included; a larger one answers an
errorand is dropped. Append audio in the small chunks it is captured in rather than whole files. - Uncommitted audio — under manual turns (
turn_detection: null), at most 5.7 MB of decoded audio may be buffered before aninput_audio_buffer.commit(about 2 minutes of 24 kHz PCM, longer for G.711); past that the append answers anerror. Commit each turn, or clear the buffer withinput_audio_buffer.clear. - Addressable items — the session keeps its 200 most recent conversation items addressable, dropping the oldest past that.
conversation.item.truncate,.retrieveand.deleteanswer anerrorfor an item that has fallen out; the model's own memory of the conversation is unaffected.
Request headers¶
The Amazon Bedrock guardrail headers apply to a realtime session. Send them on the WebSocket handshake, or on POST /v1/realtime/calls for a WebRTC call. All headers are optional.
Content Safety (Guardrails)¶
| Header | Purpose | Valid Values |
|---|---|---|
X-Amzn-Bedrock-GuardrailIdentifier | Guardrail ID for content filtering | Your guardrail identifier |
X-Amzn-Bedrock-GuardrailVersion | Guardrail version | Version number (e.g., 1) |
The guardrail selected here is the one applied per turn, as described under Guardrail coverage. Both headers are honoured only when AWS_BEDROCK_ALLOW_GUARDRAIL_OVERRIDE is enabled — otherwise the deployment's configured guardrail applies. X-Amzn-Bedrock-Trace is accepted but has no effect on this route — no guardrail trace is returned.
These headers need the deployment's own credentials
A connection opened with an ephemeral client secret is client-held, so its headers are discarded and the deployment's configured guardrail applies. Only a connection authenticated with the deployment's API key can select a guardrail per session.
No performance headers on this route
X-Amzn-Bedrock-Service-Tier and X-Amzn-Bedrock-PerformanceConfig-Latency have no effect here: a session runs on a bidirectional model stream, which carries neither a service tier nor a performance configuration.
Detailed Documentation
For complete information about these headers, configuration options, and use cases, see:
Try it¶
Python (server-side, with the official SDK)¶
Requires the openai package with realtime support (pip install "openai[realtime]"). Point the client's base_url at this deployment; client.realtime.connect(...) derives the WebSocket URL from it automatically.
import base64
from openai import AsyncOpenAI
client = AsyncOpenAI(
api_key="YOUR_API_KEY", base_url="https://your-deployment.example.com/v1"
)
async def main() -> None:
async with client.realtime.connect(model="amazon.nova-2-sonic-v1:0") as connection:
await connection.session.update(
session={
"type": "realtime",
"instructions": "You are a concise, friendly voice assistant.",
}
)
# Stream 24 kHz, 16-bit, mono, little-endian PCM in the small chunks it
# is captured in -- ~100 ms each here, well under the per-event limit.
with open("question.pcm", "rb") as audio_file:
while chunk := audio_file.read(4800):
await connection.input_audio_buffer.append(
audio=base64.b64encode(chunk).decode()
)
await connection.input_audio_buffer.commit()
await connection.response.create()
async for event in connection:
if event.type == "response.output_audio.delta":
# event.delta is base64-encoded audio in the session's output format.
...
elif event.type == "response.output_audio_transcript.delta":
print(event.delta, end="", flush=True)
elif event.type == "response.done":
break
Browser (ephemeral client secret)¶
Mint the secret from your backend — never expose the deployment's own API key to the browser — then connect directly from client-side JavaScript using the openai-insecure-api-key.<secret> subprotocol, since a browser cannot set custom WebSocket headers:
// Fetched from your own backend, which called POST /v1/realtime/client_secrets
const { value: ephemeralSecret } = await fetch("/api/realtime-secret").then((r) => r.json());
const ws = new WebSocket(
"wss://your-deployment.example.com/v1/realtime?model=amazon.nova-2-sonic-v1:0",
["realtime", `openai-insecure-api-key.${ephemeralSecret}`],
);
ws.addEventListener("open", async () => {
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
// Your own capture code: resample the microphone to the session's input
// format (24 kHz, 16-bit, mono PCM by default) and hand over small chunks.
captureAudioChunks(stream, (pcmChunk) => {
ws.send(
JSON.stringify({
type: "input_audio_buffer.append",
audio: btoa(String.fromCharCode(...new Uint8Array(pcmChunk))),
}),
);
});
// Nothing else is needed: server voice activity detection ends each turn and
// starts the answer. Under `turn_detection: null`, send
// `input_audio_buffer.commit` yourself instead.
});
ws.addEventListener("message", (event) => {
const serverEvent = JSON.parse(event.data);
if (serverEvent.type === "response.output_audio.delta") {
// serverEvent.delta is base64-encoded audio in the session's output format.
}
});
Next steps¶
Next: Search Models API · Text to Speech API · Transcriptions API · WebRTC and SIP ingress