Text to Speech API¶
Generate natural-sounding speech from text with Amazon Polly through an OpenAI-compatible interface.
At a glance¶
- Four Amazon Polly engines — Standard, Neural, Long-Form and Generative, each registered as its own model with its own voice set, see Models.
- 60+ voices across 30+ languages — name a Polly voice ID directly, or send an OpenAI voice name and let one be picked for you, see Feature compatibility.
- Language detected from the first 500 characters — Amazon Comprehend reads that much of the input to choose the voice behind an OpenAI voice name, see Working with Amazon Polly.
- SSML markup accepted — pronunciation, emphasis, pauses and prosody, up to 6,000 characters including the markup, which is not billed, see Working with Amazon Polly.
- Up to 100,000 input characters against OpenAI's 4,096 — 3,000 with no bucket configured, 20,000 on a generative voice, see Long Input.
-
instructionsis accepted and ignored,speedis rejected with SSML input — see Limits and behaviour to know.
curl -OJ -X POST "$BASE/v1/audio/speech" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "amazon.polly-neural", "voice": "Amy", "input": "Hello from Amazon Polly."}'
Endpoints¶
| Endpoint | Method | What It Does | Powered By | MCP Tool |
|---|---|---|---|---|
/v1/audio/speech | POST | Turn text into natural-sounding speech | Amazon Polly + Amazon Comprehend | openai_audio_speech |
Feature compatibility¶
| Feature | Status | Notes |
|---|---|---|
| Voice Selection | ||
| OpenAI voice names | Mapped to Polly voices | |
| Polly voice IDs | 60+ voices across 30+ languages | |
| Dynamic voice selection | Select best Polly voice based on the detected language | |
| Custom voice object | voice also accepts {"id": "Joanna"}; the id names the voice, as the plain string does | |
| Input | ||
| Plain text | Standard text input | |
| SSML markup | Fine-grained speech control | |
| Long input | Up to 100,000 characters, 24× OpenAI's limit — 20,000 with a generative voice, 100,000 with a bucket | |
| Output Formats | ||
| MP3 | Native Polly format | |
| PCM | 24 kHz per OpenAI's contract; resampled from Polly's native rate | |
| Opus | Native Polly format | |
| AAC | Encoded from PCM | |
| FLAC | Encoded from PCM | |
| WAV | Encoded from PCM | |
| OGG (Vorbis) | Native Polly format | |
| Control | ||
speed parameter | 0.2x to 4.0x playback speed, covering OpenAI's full range; generative voices speak no faster than 2.0x; rejected with SSML input (set the speed in SSML instead) | |
instructions parameter | Accepted for OpenAI API compatibility and ignored (no Amazon Polly equivalent) | |
| Extra model-specific params | Extra model-specific parameters via JSON body | |
| Streaming | ||
| Byte streaming | Default streaming mode | |
| SSE streaming | Event-based streaming | |
| Usage tracking | ||
| Input text tokens | Characters count (billing unit) | |
| Output tokens | Not available |
Legend:
- Supported — Fully compatible with OpenAI API
- Extra Feature — Enhanced capability beyond OpenAI API
- Unsupported — Not available in this implementation
Models¶
Amazon Polly Models¶
| Model | Polly Engine | Notes |
|---|---|---|
amazon.polly-standard | Standard | Lowest cost, widest language coverage |
amazon.polly-neural | Neural | Higher-quality, natural-sounding voices |
amazon.polly-long-form | Long-form | Expressive voices for narration-length content |
amazon.polly-generative | Generative | Most human-like, conversational voices; speaks long input with no bucket |
Each engine supports a different subset of voices and languages — see the Polly voice list for details. OpenAI voice names work with every model through automatic language detection and voice selection, or specify any Polly voice ID directly for 60+ voices across 30+ languages.
OpenAI Model Compatibility
stdapi.ai includes built-in model aliases that map OpenAI model names to Amazon Polly engines:
tts-1→amazon.polly-standardtts-1-hd→amazon.polly-neural
These aliases let an OpenAI-based tool reach these models under the names it already sends, with no configuration change. You can also customize or override these aliases to suit your needs.
Working with Amazon Polly¶
Amazon Polly Features¶
- SSML Support : Fine-grained control over pronunciation, emphasis, pauses, and prosody — SSML docs
- Flexible Formats: mp3, ogg, wav, flac, aac, opus, pcm
- Streaming Options: Raw bytes (default) or SSE events with
stream_format: "sse" - Speed Control: Adjust playback from 0.2x to 4.0x — generative voices speak no faster than 2.0x
- Speech Marks: Word, sentence, viseme, and SSML timing metadata with
SpeechMarkTypes(returned as JSON instead of audio)
Performance Tips: Optimize Speed & Cost
- Prefer mp3 or ogg for the lowest latency;
pcmis returned at OpenAI's 24 kHz contract unless you request an explicitSampleRate(see Sample Rate) - Specify a Polly voice ID to bypass language detection—faster responses, no Amazon Comprehend charges
- Configure a default language via
DEFAULT_TTS_LANGUAGEenvironment variable to skip language detection for all requests using OpenAI voice names
Language Detection Behavior
When using OpenAI voice names without specifying a default language, the system analyzes only the first 500 characters of your text to detect the language. This approach:
- Works best with long, single-language texts where the first 500 characters are representative
- May be inconsistent with very short texts (< 100 characters) where language detection has limited context
- Can produce mixed results with multi-language content where different parts use different languages
For consistent behavior across requests, consider:
- Setting
DEFAULT_TTS_LANGUAGEfor applications serving primarily one language - Using Polly voice IDs directly when you know the target language
- Structuring multi-language applications to make separate API calls per language
Default Streaming Mode: API vs MCP
- API usage: Default is byte streaming (raw audio data)
- MCP tool usage: Default is SSE streaming (
stream_format: "sse")
When used as an MCP tool, the response defaults to SSE events (speech.audio.delta, speech.audio.done) for better client compatibility. Override by explicitly setting stream_format: "audio" in your request.
Long Input¶
A single request accepts up to 3,000 characters (6,000 including SSML markup, which is not billed). Beyond that, how the audio is produced depends on the voice.
Generative voices (amazon.polly-generative) speak up to 20,000 characters without any further configuration, and the audio starts arriving while the rest is still being spoken. Two cases keep the behaviour described below instead: an input written as an SSML document, and a request using SpeechMarkTypes.
Every other voice, and generative input longer than 20,000 characters, is synthesized into an S3 bucket co-located with the serving region, so longer input requires a bucket in a region that can serve the request: a AWS_S3_REGIONAL_BUCKETS entry for that region, or AWS_S3_BUCKET when the region is the first AWS_BEDROCK_REGIONS entry — the only region AWS_S3_BUCKET covers. With AWS_POLLY_REGION pinned to any other region, a regional-bucket entry is the only option. The audio object is written under AWS_S3_TMP_PREFIX and deleted once the request ends. On a request that ends before Amazon Polly has finished — a timeout, a failure, or a client that disconnected — the deletion is issued while the synthesis is still running, so an object written after it is not removed by the request: the recommended lifecycle rule on that prefix is what expires it, and it must be in place.
- With a bucket configured: up to 100,000 characters (200,000 including SSML markup), against OpenAI's 4,096-character limit. Expect roughly one extra second per 1,000 characters, bounded by
AI_RESPONSE_TIMEOUT. - Without one: a generative voice still speaks up to 20,000 characters; every other request above 3,000 characters is rejected, however long, with the length the server does accept — so callers can split their text.
Same response either way
Nothing else changes: the response is the complete audio file in the requested response_format, and stream_format: "sse" still delivers speech.audio.delta events.
A configured guardrail caps the input first
When a guardrail applies to the request, the whole input is checked in a single Amazon Bedrock ApplyGuardrail call, which has a maximum input size of its own: a Service Quotas value per guardrail policy, counted in text units of 1,000 characters, that differs between AWS Regions — as low as 25 text units (25,000 characters) in some, 1,000 in others.
The reachable input length is therefore the smaller of the two: the limit above, and the quota in the guardrail's Region. Beyond the quota the request fails with 429 before any audio is synthesized, so raise the maximum input size quotas for the policies your guardrail applies, or keep requests under them.
Long input is billed on acceptance
Amazon Polly bills the whole input as soon as it accepts a long request, before the audio is produced. A request that then reaches AI_RESPONSE_TIMEOUT, fails, or is abandoned by the client is charged in full and still counted in the request usage and cost records — retrying it pays for the text twice.
Provider-Specific Parameters¶
Unlock advanced Amazon Polly capabilities by passing provider-specific parameters directly in your requests. These parameters are forwarded to Polly's SynthesizeSpeech API and allow you to access features unique to Polly.
How It Works:
Add provider-specific fields at the top level of your request body alongside standard OpenAI parameters. The API automatically forwards these to Amazon Polly.
Examples:
Lexicon Support:
Apply custom pronunciation lexicons to your speech synthesis:
{
"model": "amazon.polly-neural",
"voice": "Joanna",
"input": "Amazon Polly uses lexicons for custom pronunciation.",
"response_format": "mp3",
"LexiconNames": ["MyCustomLexicon"]
}
Sample Rate:
Specify custom audio sample rate (8000, 16000, 22050, or 24000 Hz; PCM output supports 8000 and 16000 only):
{
"model": "amazon.polly-neural",
"voice": "Matthew",
"input": "High quality audio at 24kHz.",
"response_format": "mp3",
"SampleRate": "24000"
}
PCM output defaults to OpenAI's 24 kHz contract
Per OpenAI's TTS API, response_format: "pcm" is raw, headerless 24 kHz 16-bit mono little-endian audio. Without an explicit SampleRate, pcm output is synthesized by Polly at 16 kHz and resampled to 24 kHz server-side. Pass a Polly-native SampleRate (8000 or 16000) to skip resampling and receive Polly's raw rate instead — since raw PCM carries no embedded rate, only do this when your client knows to play it back at that rate.
Language Code:
Specify the language for bilingual voices (only useful for voices that support multiple languages):
{
"model": "amazon.polly-neural",
"voice": "Aditi",
"input": "Hello, how are you?",
"response_format": "mp3",
"LanguageCode": "en-IN"
}
Speech Marks:
Request word, sentence, viseme, or SSML timing marks instead of audio (useful for lip-sync, karaoke-style highlighting, or subtitle alignment):
{
"model": "amazon.polly-neural",
"voice": "Joanna",
"input": "Hello, how are you?",
"SpeechMarkTypes": ["word", "sentence"]
}
Speech marks return JSON, not audio
When SpeechMarkTypes is set, Polly returns timing metadata only. The response is a stream of JSON objects (one per line) with the application/x-json-stream content type:
response_formatis ignored — no audio is returned.stream_format: "sse"is rejected with HTTP 400, since the payload is not audio events.- The
ssmlmark type requires SSML input (<speak>…</speak>); requesting it with plain text returns HTTP 400.
{"time":0,"type":"word","start":0,"end":5,"value":"Hello"}
{"time":576,"type":"word","start":7,"end":10,"value":"how"}
Configuration Options:
Option 1: Per-Request
Add provider-specific parameters directly in your request body (as shown in examples above).
Option 2: Server-Wide Defaults
Configure default parameters for specific models via the DEFAULT_MODEL_PARAMS environment variable:
export DEFAULT_MODEL_PARAMS='{
"amazon.polly-neural": {
"SampleRate": "24000"
}
}'
Note: Per-request parameters override server-wide defaults.
Behavior:
Compatible parameters are forwarded to Polly and applied; unsupported parameters return HTTP 400 with an error message.
Available Parameters:
The following parameters from the Amazon Polly SynthesizeSpeech API can be used:
LexiconNames(list): Apply pronunciation lexiconsSampleRate(string): Audio sample rate in Hz —8000,16000,22050, or24000(pcmoutput:8000or16000; omit it to get OpenAI's 24 kHzpcmcontract instead)LanguageCode(string): Language code for bilingual voices only (e.g.,en-IN,hi-IN)SpeechMarkTypes(list): Timing marks to return instead of audio —sentence,ssml,viseme,word
Limits and behaviour to know¶
instructionsis accepted and ignored. Amazon Polly has no equivalent parameter, so the audio comes back as if the field had not been sent — use SSML, or a different voice, to get the delivery you want.speedis rejected with SSML input, because SSML carries a speaking rate of its own: set it with<prosody>instead.speedruns from0.2to4.0. Above2.0a generative voice (amazon.polly-generative) speaks no faster: the request succeeds and the audio is simply no shorter. Every other voice speeds up across the whole range.- Usage is counted in characters, the native billing unit of Amazon Polly and Amazon Comprehend, rather than in OpenAI-style tokens. No output token count is reported.
- Once an SSE stream is accepted, a synthesis failure can no longer be an HTTP status: it arrives as a terminal
errorevent inside the200response, andspeech.audio.doneis then omitted. - Input above 3,000 characters needs an S3 bucket unless the voice is generative, and a long request is billed the moment Amazon Polly accepts it — both are detailed under Long Input.
Request headers¶
This endpoint supports standard Bedrock headers for enhanced control over your requests. All headers are optional and can be combined as needed.
Content Safety (Guardrails)¶
| Header | Purpose | Valid Values |
|---|---|---|
X-Amzn-Bedrock-GuardrailIdentifier | Guardrail ID for content filtering | Your guardrail identifier |
X-Amzn-Bedrock-GuardrailVersion | Guardrail version | Version number (e.g., 1) |
The guardrail evaluates the text to synthesize; the audio produced from it is not itself evaluated. X-Amzn-Bedrock-Trace is accepted but has no effect on this route — no guardrail trace is returned.
Example with headers:
curl -X POST "$BASE/v1/audio/speech" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-H "X-Amzn-Bedrock-GuardrailIdentifier: your-guardrail-id" \
-H "X-Amzn-Bedrock-GuardrailVersion: 1" \
-d '{
"model": "amazon.polly-neural",
"voice": "Amy",
"input": "Welcome to the future of voice technology!"
}' \
--output speech.mp3
No performance headers on this route
X-Amzn-Bedrock-Service-Tier and X-Amzn-Bedrock-PerformanceConfig-Latency have no effect here: speech is synthesized by Amazon Polly, which is not invoked through the Amazon Bedrock runtime.
Detailed Documentation
For complete information about these headers, configuration options, and use cases, see:
Try it¶
Stream audio as bytes (default):
curl -OJ -X POST "$BASE/v1/audio/speech" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "amazon.polly-neural",
"voice": "Amy",
"input": "Welcome to the future of voice technology!",
"response_format": "mp3"
}'
Stream audio as SSE events:
curl -N -X POST "$BASE/v1/audio/speech" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "amazon.polly-neural",
"voice": "Amy",
"input": "This audio streams as SSE events!",
"response_format": "mp3",
"stream_format": "sse"
}'
Next steps¶
Next: Models API · Speech to Text API · Model aliases for tts-1 and tts-1-hd · Text-to-speech IAM permissions, including long input