Speech to Text API¶
Transcribe audio to text with Amazon Transcribe or Amazon Bedrock audio-capable models through an OpenAI-compatible interface.
At a glance¶
- Amazon Transcribe across 100+ languages, or any Bedrock model that accepts the
SPEECHmodality — including Amazon Nova Sonic and Mistral Voxtral, see Models. - Each phrase sent as it is recognized —
stream=truestreams phrase by phrase whenever the request names the language to expect, and that path needs no S3 bucket, see Streaming. - SRT and VTT cut by Amazon Transcribe itself — with the timings it produced, alongside
json,text,verbose_jsonanddiarized_json, see Feature compatibility. - Speaker labels
A,B, … for 10 speakers by default and 30 at most —diarized_json, in one response or streamed astranscript.text.segmentevents, see Working with Amazon Transcribe. - Medical dictation and clinical conversations —
amazon.transcribe-medical, in US English, with six specialties when streamed, see Medical transcription. - Audio as base64, data URI, HTTPS URL, S3 URI or a
file-id:reference — a JSON body alternative to the multipart upload, for MCP clients and AI agents, see Try it. -
known_speaker_namesandknown_speaker_referencesare accepted and ignored, andamazon.transcriberefusesprompt,temperature,keywordsandinclude— see Limits and behaviour to know.
curl -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file=@meeting-recording.mp3 \
-F model=amazon.transcribe
Endpoints¶
| Endpoint | Method | What It Does | Powered By | MCP Tool |
|---|---|---|---|---|
/v1/audio/transcriptions | POST | Convert spoken audio to written text | Amazon Transcribe or Amazon Bedrock Audio Models | openai_audio_transcription |
Feature compatibility¶
| Feature | Status | Notes |
|---|---|---|
| Input | ||
| Audio file upload | Multipart file upload | |
| JSON body input | Base64, data URI, HTTPS URL, S3 URI, or file-id: reference — for MCP / AI agents | |
| Output Formats | ||
json | Structured transcription | |
text | Plain text output | |
verbose_json | With timestamps and details (Amazon Transcribe; not Bedrock models); each segment's seek, temperature and tokens are placeholders — see Limits and behaviour to know | |
diarized_json | With speaker identification (Amazon Transcribe; not Bedrock models); streams as transcript.text.segment events with stream=true | |
srt | Subtitle format with timing (amazon.transcribe; not amazon.transcribe-medical nor Bedrock models); rejected with stream=true | |
vtt | WebVTT subtitle format (amazon.transcribe; not amazon.transcribe-medical nor Bedrock models); rejected with stream=true | |
| Language | ||
| Language specification | ISO-639-1 language codes | |
languages (expected languages) | Expected-language list (ISO-639-1) for multi-language audio; cannot be combined with language. Drives Amazon Transcribe multi-language identification; folded into the transcription context on Bedrock models | |
| Auto language detection | Automatic identification | |
Detected languages in response | The json response reports the detected language(s) as a languages array (Amazon Transcribe only) | |
| Streaming | ||
stream (SSE streaming) | Set stream: true to receive incremental results as server-sent events. srt and vtt are rejected rather than answered without their cues; verbose_json is accepted but degrades to text-only events — timestamps and segments are dropped | |
transcript.text.segment events | Speaker segments streamed as they are finalized, with response_format=diarized_json on a model that labels speakers | |
| Advanced | ||
timestamp_granularities | Word or segment level; requires response_format=verbose_json (Amazon Transcribe only) | |
| Speaker diarization | Automatic speaker separation; requires response_format=diarized_json (Amazon Transcribe only) | |
known_speaker_names | Accepted but ignored — diarization falls back to generic speaker labels | |
known_speaker_references | Accepted but ignored — diarization falls back to generic speaker labels | |
chunking_strategy | Only auto is accepted; other values are rejected | |
temperature | Bedrock models only; rejected by Amazon Transcribe | |
prompt | Bedrock models only; rejected by Amazon Transcribe | |
keywords | Bedrock models only (folded into the transcription context); rejected by Amazon Transcribe — use a pre-created custom vocabulary via the VocabularyName extra parameter | |
include (logprobs) | Accepted on Bedrock models but never populated (logprobs is always null); rejected by Amazon Transcribe | |
| Extra model-specific params | Amazon Transcribe optional settings via JSON body (see below) | |
| Usage tracking | ||
| Input audio duration | Seconds (billing unit on Amazon Transcribe) | |
| Output text tokens | On models from Bedrock |
Legend:
- Supported — Fully compatible with OpenAI API
- Available on Select Models — Check your model's capabilities
- Partial — Supported with limitations
- Unsupported — Not available in this implementation
- Extra Feature — Enhanced capability beyond OpenAI API
Models¶
Amazon Models¶
| Model | Supported Languages | Notes |
|---|---|---|
| amazon.transcribe | 100+ | Full-featured transcription with speaker diarization and subtitle generation at the cost of higher latency |
| amazon.transcribe-medical | US English | Clinical dictation and patient–clinician conversations, with medical terminology — see Medical transcription |
Configuration Required
You must configure a bucket to use these models, through AWS_S3_BUCKET, AWS_TRANSCRIBE_S3_BUCKET, or an AWS_S3_REGIONAL_BUCKETS entry for a region where Amazon Transcribe is a candidate. This bucket is used for temporary storage during transcription processing.
Mistral Models¶
| Model | Supported Languages | Notes |
|---|---|---|
| mistral.voxtral-mini-3b-2507 | Multilingual (auto-detected) | Compact model for fast transcription |
| mistral.voxtral-small-24b-2507 | Multilingual (auto-detected) | Larger model for enhanced accuracy |
Mistral Voxtral Limitations
Mistral Voxtral models have the following restrictions when running on Amazon Bedrock:
- File size limit: ~2MB maximum input file size
- Audio channels: Mono channel audio only (single channel)
Amazon Nova Sonic¶
| Model | Supported Languages | Notes |
|---|---|---|
| amazon.nova-2-sonic-v1:0 | Multilingual (auto-detected) | Low-cost real-time speech recognition |
Name this model to transcribe through Amazon Nova Sonic instead of Amazon Transcribe. It is the cheapest transcription option (about $0.006 per minute of audio at current Amazon Bedrock rates) and returns punctuated text in the language that was spoken. Transcription is model selection, never automatic: requests that do not name this model are unaffected.
What this model does not provide
- Response formats:
jsonandtextonly.srt,vtt,verbose_jsonanddiarized_jsonare rejected, as istimestamp_granularities— this model returns no timestamps and does not report which language it detected. Useamazon.transcribefor subtitles, timestamps, speaker diarization or detected-language reporting. - Audio length: up to 10 minutes per request. Longer recordings are rejected; use
amazon.transcribe, which has no such limit.
Other Amazon Bedrock Models¶
Any Amazon Bedrock model that accepts the SPEECH input modality through the Converse API can transcribe out of the box: the gateway sends the audio together with a transcription prompt and returns the model's text output.
Audio Input Formats on Bedrock Models
Uploads in the formats the Bedrock Converse audio block accepts — aac, flac, m4a, mka, mkv, mp3, mp4, mpeg, mpga, ogg, opus, pcm, wav, webm, and x-aac — are sent through as-is. Any other audio or video upload is automatically converted to FLAC before transcription (requires FFmpeg on the server), including the audio track of a video container. An upload that is neither audio nor video is rejected with the list of accepted formats; an audio or video file whose track cannot be decoded is rejected as carrying no decodable audio.
Working with Amazon Transcribe¶
Amazon Transcribe Features¶
Model & Features:
- Use
amazon.transcribewith the same interface as OpenAI's Whisper API - Or use OpenAI model names directly:
whisper-1,gpt-transcribe,gpt-live-transcribe,gpt-4o-transcribe, andgpt-4o-mini-transcribework out of the box (they map toamazon.transcribe) - Auto-detect language or specify it for faster processing, or list the expected languages with
languagesfor multi-language audio - Word-level or segment-level timestamps with
verbose_json - Speaker Diarization : Automatically identify and label different speakers with
diarized_json, in one response or streamed segment by segment. Each speaker gets one of the sequential capital lettersA,B, ... in the order the speakers are first heard — rolling over toAA,AB, ... past 26 speakers — and keeps it for the whole transcript - Native Subtitles : SRT/VTT files generated directly by Amazon Transcribe with precise timing
OpenAI Model Compatibility
stdapi.ai includes built-in model aliases that map the OpenAI model names to Amazon Transcribe:
whisper-1→amazon.transcribegpt-transcribe→amazon.transcribegpt-live-transcribe→amazon.transcribegpt-4o-transcribe→amazon.transcribegpt-4o-mini-transcribe→amazon.transcribe
These aliases let an OpenAI-based tool reach these models under the names it already sends, with no configuration change. You can also customize or override these aliases to suit your needs.
Performance Tips: Optimize Speed & Cost
- Specify the language if you know it—skips auto-detection for faster processing and lower AWS costs
Provider-Specific Parameters¶
Unlock advanced Amazon Transcribe capabilities by passing provider-specific parameters directly in your request body. These parameters are forwarded to Transcribe's StartTranscriptionJob API.
JSON body required
Unlike the multipart/form-data upload, extra parameters are only reachable through the application/json request body (file as base64, data URI, HTTPS URL, or file-id: reference) — the multipart path only accepts the documented OpenAI fields.
PII Redaction:
Redact personally identifiable information from the transcript (only the single-output redacted mode is supported):
{
"model": "amazon.transcribe",
"file": "data:audio/mp3;base64,<base64-encoded-audio>",
"ContentRedaction": {
"RedactionType": "PII",
"PiiEntityTypes": ["NAME", "SSN", "CREDIT_DEBIT_NUMBER"]
}
}
Custom Vocabulary and Filtering:
Improve recognition of domain-specific terms and mask or remove profanity/sensitive words:
{
"model": "amazon.transcribe",
"file": "data:audio/mp3;base64,<base64-encoded-audio>",
"VocabularyName": "MyCustomVocabulary",
"VocabularyFilterName": "MyProfanityFilter",
"VocabularyFilterMethod": "mask"
}
Alternative Transcriptions and Channel Identification:
Request multiple candidate transcriptions per segment, or transcribe each audio channel separately (e.g. two-party phone calls recorded in stereo):
{
"model": "amazon.transcribe",
"file": "data:audio/mp3;base64,<base64-encoded-audio>",
"ShowAlternatives": true,
"MaxAlternatives": 3,
"ChannelIdentification": true
}
Incompatible with diarized_json
ChannelIdentification cannot be combined with response_format=diarized_json, which already forces AWS speaker-label diarization. Requesting both returns HTTP 400.
Toxicity Detection:
Flag toxic content (profanity, hate speech, harassment) in the transcript:
{
"model": "amazon.transcribe",
"file": "data:audio/mp3;base64,<base64-encoded-audio>",
"ToxicityDetection": [{"ToxicityCategories": ["ALL"]}]
}
Multi-Language Identification:
Detect and transcribe multiple languages spoken in the same audio, optionally restricted to a candidate list:
{
"model": "amazon.transcribe",
"file": "data:audio/mp3;base64,<base64-encoded-audio>",
"IdentifyMultipleLanguages": true,
"LanguageOptions": ["en-US", "es-US", "fr-FR"]
}
Per-Language Custom Resources:
A custom vocabulary, vocabulary filter or custom language model can only be attached per candidate language when the language is identified rather than given. Pair LanguageIdSettings with LanguageOptions so the dialect your resources were created for is the one identified:
{
"model": "amazon.transcribe",
"file": "data:audio/mp3;base64,<base64-encoded-audio>",
"LanguageOptions": ["en-US", "es-US"],
"LanguageIdSettings": {
"en-US": {"VocabularyName": "MedicalTermsEnUs"},
"es-US": {"VocabularyName": "MedicalTermsEsUs"}
}
}
Identification required
LanguageIdSettings applies to identified languages only. Combined with a fixed language (or a single-entry languages), the request returns HTTP 400 rather than silently dropping the custom resources — use the flat VocabularyName, VocabularyFilterName and ModelSettings parameters in that case. LanguageModelName is not available with IdentifyMultipleLanguages.
Standard languages parameter
The standard OpenAI languages parameter drives the same multi-language identification with plain ISO-639-1 codes (e.g. ["en", "es", "fr"]), works on the multipart path too, and the detected language(s) come back in the json response's languages array. A single-entry list behaves like language. Do not combine it with language or with the provider-specific parameters above.
Configuration Options:
Option 1: Per-Request
Add provider-specific parameters directly in your JSON request body (as shown in examples above).
Option 2: Server-Wide Defaults
Configure default parameters for amazon.transcribe via the DEFAULT_MODEL_PARAMS environment variable:
export DEFAULT_MODEL_PARAMS='{
"amazon.transcribe": {
"VocabularyFilterName": "MyProfanityFilter",
"VocabularyFilterMethod": "mask"
}
}'
Note: Per-request parameters override server-wide defaults.
Behavior:
Compatible parameters are forwarded to Amazon Transcribe and applied; unsupported parameters or values return HTTP 400 with an error message.
Available Parameters:
The following parameters from Amazon Transcribe's StartTranscriptionJob API can be used:
ContentRedaction(object): PII redaction —RedactionType(PII),PiiEntityTypes(list),RedactionOutput(redactedonly;redacted_and_unredactedis rejected — the unredacted copy is not tracked for automatic cleanup)VocabularyName(string): Custom vocabulary to improve recognition accuracyVocabularyFilterName/VocabularyFilterMethod(string /mask,remove,tag): Profanity or sensitive-word filteringShowAlternatives/MaxAlternatives(bool / integer2-10): Return multiple candidate transcriptions per segmentChannelIdentification(bool): Transcribe each audio channel separately (incompatible withdiarized_json)MaxSpeakerLabels(integer2-30): Maximum speakers to identify withresponse_format=diarized_json(default10). It has no effect on a streamed transcription delivered phrase by phrase (see Streaming). Amazon Transcribe distinguishes at most 30 speakers whatever the value, and separates them less reliably past about fiveShowSpeakerLabels(bool): Always on withresponse_format=diarized_json; setting it directly with another format runs AWS speaker labeling without exposing speaker data in the responseToxicityDetection(list): Toxic-content flagging —[{"ToxicityCategories": ["ALL"]}]IdentifyMultipleLanguages/LanguageOptions(bool / list): Multi-language identification, optionally restricted to a candidate list (supersedeslanguage; cannot be combined with the standardlanguagesparameter)LanguageIdSettings(object): Per-languageVocabularyName,VocabularyFilterNameandLanguageModelName, keyed by language code (up to five) — the only way to attach them when the language is identified rather than givenModelSettings(object):LanguageModelName— custom language model selection
VocabularyName, VocabularyFilterName, and custom language models must already exist in your AWS account (created via the AWS Transcribe console, CLI, or SDK) before being referenced here.
Medical transcription¶
amazon.transcribe-medical transcribes clinical audio — a physician's dictated notes, a patient–clinician conversation — with Amazon Transcribe Medical, which recognizes medical terms, drug names and dosages. It is priced separately from amazon.transcribe (see Amazon Transcribe pricing) and needs its own IAM permissions.
{
"model": "amazon.transcribe-medical",
"file": "data:audio/wav;base64,<base64-encoded-audio>",
"Type": "DICTATION"
}
| Parameter | Values | Notes |
|---|---|---|
Type | CONVERSATION (default), DICTATION | One speaker dictating, or a conversation between several |
Specialty | PRIMARYCARE (default), CARDIOLOGY, NEUROLOGY, ONCOLOGY, RADIOLOGY, UROLOGY | Any value other than PRIMARYCARE requires stream=true (without it, HTTP 400) and a deployment serving live medical transcription (without it, HTTP 503) |
VocabularyName | string | A medical custom vocabulary you created |
ShowSpeakerLabels, MaxSpeakerLabels, ChannelIdentification, ShowAlternatives, MaxAlternatives | as for amazon.transcribe | Not carried by a phrase-by-phrase stream, so combining them with Specialty returns HTTP 400 |
- Language: US English only.
languageandlanguagesmay be omitted or set toen; any other language returns HTTP 400, as does anyamazon.transcribeparameter this model does not list above (redaction, toxicity detection, language identification, vocabulary filters, custom language models,ContentIdentificationType). - Formats:
json,text,verbose_jsonanddiarized_json.srtandvttreturn HTTP 400: requestverbose_jsonfor timed segments. - Streaming:
stream=truestreams phrase by phrase, since the language is always known. A request setting a parameter a stream cannot carry, or reaching a deployment where no configured region offers live medical transcription, is still served, all at once when the transcript is complete — except with aSpecialtyother thanPRIMARYCARE, which needs the live path (HTTP 503 without it). - Translation: not available — the transcript is already in English.
- Regions: medical transcription is offered in fewer regions than
amazon.transcribe, and fewer still when streamed (supported regions). A request moves on to the next configured region that offers it; when none does, it returns HTTP 503. -
Fixed variants: a client that can only send a model name (Home Assistant, Open WebUI) reaches a specific audio
Typethrough an alias that carries configuration. ASpecialtyother thanPRIMARYCAREin an alias only works for clients that sendstream=true; the others get HTTP 400.export MODEL_ALIASES='{ "medical-dictation": { "model": "amazon.transcribe-medical", "extra_params": {"Type": "DICTATION"} } }'
Protected health information
Amazon Transcribe Medical is a HIPAA-eligible service; processing protected health information requires a Business Associate Addendum with AWS and a deployment configured for it. A request served as a transcription job — every non-streamed request, and a streamed one that falls back to a job — writes the audio and the transcript under the AWS_S3_TMP_PREFIX of your transcription bucket and deletes them after responding, best effort; a lifecycle rule expiring that prefix removes whatever an interrupted server leaves behind. In a versioned bucket a deletion keeps a noncurrent version: the buckets the Terraform module creates expire both current and noncurrent versions under that prefix after one day, so the data can remain for a day or two; a bucket you bring needs the same noncurrent-version expiration on that prefix, or deleted audio and transcripts stay recoverable indefinitely. Transcripts also reach the server log when LOG_REQUEST_PARAMS is enabled — keep it off for this model.
Streaming¶
stream=true returns the transcript as server-sent events — transcript.text.delta events followed by a final transcript.text.done — instead of one response body. A ready-to-run example is below.
Each phrase is sent as it is recognized, rather than after the whole recording, whenever the request names the language to expect: send language, or two or more expected languages. That path needs no S3 bucket, so a deployment with no storage configured serves streamed transcriptions.
When no configured region can open a live session — the deployment lacks the transcribe:StartStreamTranscription permission (transcribe:StartMedicalStreamTranscription for the medical model), or no region offers streaming for the model — the request is served as a transcription job instead: it is still streamed, but its events arrive together at the end, and it needs a transcription bucket.
A request naming neither is still streamed, but its events arrive together once the recording has been read and its language detected. Operators can set AWS_TRANSCRIBE_STREAM_LANGUAGES to the languages their callers actually send, which gives those requests the faster path too. The same applies to a request using any provider-specific parameter above other than VocabularyName, VocabularyFilterName and VocabularyFilterMethod, which are the only ones a phrase-by-phrase transcript can carry.
Speaker segments¶
With response_format=diarized_json, each stretch of transcript spoken by one speaker is also sent as a transcript.text.segment event, carrying the segment's id, start, end, speaker and text. Segments are interleaved with the deltas and always precede the final transcript.text.done.
A speaker is attached only once the recognizer has settled the words it labels, so a segment is sent once and its speaker never changes afterwards. Segments carry the same speaker labels as the non-streamed diarized_json response. Naming the language is what gets the segments phrase by phrase instead of together at the end: send language, or two or more expected languages.
Each delta belonging to a segment names that segment's id in its segment_id field, so the two can be correlated without matching their text. Text no speaker was attributed to — including text the recognizer never settled — arrives as a delta with no segment_id and no segment event of its own, so the concatenated deltas — not the segments — are the complete transcript.
A model that reports no speakers rejects diarized_json with stream=true rather than answering with an unlabelled transcript.
Streamed events carry no subtitle cues, which is why srt and vtt are rejected with stream=true.
Limits and behaviour to know¶
- With
amazon.transcribe, theprompt,temperature,keywordsandincludeparameters are rejected with an error to ensure consistent transcription accuracy. Forkeywords, the error points at the pre-created custom vocabulary alternative via theVocabularyNameextra parameter. known_speaker_namesandknown_speaker_referencesare accepted but ignored for every model: Amazon Transcribe's automatic speaker diarization runs without known speaker references, so it falls back to generic speaker labels.include: ["logprobs"]is accepted on Bedrock models but never populated —logprobscomes backnull— because the Converse API returns no token log probabilities.- With
amazon.transcribe, the decoder fields of averbose_jsonsegment are placeholders:seekis always0,temperaturealways0.0andtokensalways empty, whileavg_logprobandno_speech_probreport the documented silence signal (-2.0and1.0on a segment carrying no speech,0.0and0.0otherwise) rather than measured probabilities. The model reports no per-segment decoder state, sostart,endandtextare the values that carry information: a transcript cannot be re-aligned or re-decoded from the rest. chunking_strategyacceptsautoonly; any other value is rejected rather than silently applied.amazon.transcribe-medicalaccepts US English only, refusessrtandvtt, and takes a specialty other than primary care only withstream=trueon a deployment serving live medical transcription — see Medical transcription.- The extra Amazon Transcribe parameters are reachable through the
application/jsonbody only, not through the multipart upload — see Provider-Specific Parameters.
Request headers¶
This endpoint supports standard Bedrock headers for enhanced control over your requests. All headers are optional and can be combined as needed.
Content Safety (Guardrails)¶
| Header | Purpose | Valid Values |
|---|---|---|
X-Amzn-Bedrock-GuardrailIdentifier | Guardrail ID for content filtering | Your guardrail identifier |
X-Amzn-Bedrock-GuardrailVersion | Guardrail version | Version number (e.g., 1) |
The guardrail evaluates the transcript the model produced, not the audio sent. On a streamed request the events are withheld until the transcript is complete, so the guardrail sees the whole text before any of it is delivered. When it masks content, the stream carries the masked transcript as a single delta and the speaker segments are withheld, each one repeating text the guardrail took out — so a streamed diarized_json request answers with the masked transcript, where the same request without stream=true fails with content_filter instead. X-Amzn-Bedrock-Trace is accepted but has no effect on this route — no guardrail trace is returned.
Performance Optimization¶
| Header | Purpose | Valid Values |
|---|---|---|
X-Amzn-Bedrock-Service-Tier | Service tier selection | default, flex, priority, reserved |
X-Amzn-Bedrock-PerformanceConfig-Latency | Latency optimization | standard, optimized |
Both apply only to models transcribed through the Amazon Bedrock runtime. They have no effect on amazon.transcribe, which is served by Amazon Transcribe, nor on Amazon Nova Sonic, which is served over its own bidirectional stream.
Example with headers:
curl -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "X-Amzn-Bedrock-GuardrailIdentifier: your-guardrail-id" \
-H "X-Amzn-Bedrock-GuardrailVersion: 1" \
-F file=@meeting-recording.mp3 \
-F model=amazon.transcribe \
-F response_format=json
Detailed Documentation
For complete information about these headers, configuration options, and use cases, see:
Try it¶
Transcribe audio to JSON:
curl -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file=@meeting-recording.mp3 \
-F model=amazon.transcribe \
-F response_format=json
Transcribe via JSON body (MCP and AI agents):
When using MCP tools or HTTP clients that cannot construct multipart requests, pass the audio as a data URI or URL:
# Data URI (inline base64)
curl -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"file": "data:audio/mp3;base64,<base64-encoded-audio>",
"model": "amazon.transcribe",
"response_format": "json"
}'
# HTTPS URL (server fetches the audio)
curl -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"file": "https://example.com/audio.mp3",
"model": "amazon.transcribe"
}'
# Files API reference (file-id: URI scheme)
curl -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"file": "file-id:file-0190c51c7de7455d9b8c2efe27dfbf67",
"model": "amazon.transcribe"
}'
See Files API → Referencing Uploaded Files for the full description of the file-id: URI scheme.
Generate subtitles:
curl -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file=@video-audio.mp3 \
-F model=amazon.transcribe \
-F response_format=srt \
-F language=en
Transcribe with speaker diarization:
curl -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file=@meeting-recording.mp3 \
-F model=amazon.transcribe \
-F response_format=diarized_json
Stream a transcription as SSE events:
Set stream=true to receive the transcript incrementally as server-sent events (transcript.text.delta events followed by a final transcript.text.done event):
curl -N -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file=@meeting-recording.mp3 \
-F model=amazon.transcribe \
-F language=en \
-F stream=true
Stream speaker segments as they are recognized:
curl -N -X POST "$BASE/v1/audio/transcriptions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F file=@meeting-recording.mp3 \
-F model=amazon.transcribe \
-F response_format=diarized_json \
-F language=en \
-F stream=true
Each completed speaker segment arrives as a transcript.text.segment event beside the text deltas:
data: {"delta":"Good morning, how can I help?","type":"transcript.text.delta","segment_id":"seg_0"}
data: {"id":"seg_0","start":0.01,"end":4.27,"speaker":"A","text":"Good morning, how can I help?","type":"transcript.text.segment"}
Naming the language starts the transcript sooner
language is what lets each phrase be sent as it is recognized rather than after the whole recording — see Streaming Transcriptions.
verbose_json streams as plain text
stream=true combined with response_format=verbose_json is accepted rather than rejected, but the streamed events carry transcript.text.delta / .done only — segment timings, word timings and language details are not included. Request verbose_json without stream to get them.
Next steps¶
Next: Models API · Speech to English API · AWS_TRANSCRIBE_S3_BUCKET, where non-streamed jobs are staged · Speech-to-text IAM permissions