Skip to content

Speech to Text API

Transcribe audio to text with Amazon Transcribe or Amazon Bedrock audio-capable models through an OpenAI-compatible interface.

At a glance

  • Amazon Transcribe across 100+ languages, or any Bedrock model that accepts the SPEECH modality — including Amazon Nova Sonic and Mistral Voxtral, see Models.
  • Each phrase sent as it is recognized — stream=true streams phrase by phrase whenever the request names the language to expect, and that path needs no S3 bucket, see Streaming.
  • SRT and VTT cut by Amazon Transcribe itself — with the timings it produced, alongside json, text, verbose_json and diarized_json, see Feature compatibility.
  • Speaker labels A, B, … for 10 speakers by default and 30 at most — diarized_json, in one response or streamed as transcript.text.segment events, see Working with Amazon Transcribe.
  • Medical dictation and clinical conversations — amazon.transcribe-medical, in US English, with six specialties when streamed, see Medical transcription.
  • Audio as base64, data URI, HTTPS URL, S3 URI or a file-id: reference — a JSON body alternative to the multipart upload, for MCP clients and AI agents, see Try it.
  • known_speaker_names and known_speaker_references are accepted and ignored, and amazon.transcribe refuses prompt, temperature, keywords and include — see Limits and behaviour to know.
curl -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F file=@meeting-recording.mp3 \
  -F model=amazon.transcribe

Endpoints

Endpoint Method What It Does Powered By MCP Tool
/v1/audio/transcriptions POST Convert spoken audio to written text Amazon Transcribe or Amazon Bedrock Audio Models openai_audio_transcription

Feature compatibility

Feature Status Notes
Input
Audio file upload Multipart file upload
JSON body input Base64, data URI, HTTPS URL, S3 URI, or file-id: reference — for MCP / AI agents
Output Formats
json Structured transcription
text Plain text output
verbose_json With timestamps and details (Amazon Transcribe; not Bedrock models); each segment's seek, temperature and tokens are placeholders — see Limits and behaviour to know
diarized_json With speaker identification (Amazon Transcribe; not Bedrock models); streams as transcript.text.segment events with stream=true
srt Subtitle format with timing (amazon.transcribe; not amazon.transcribe-medical nor Bedrock models); rejected with stream=true
vtt WebVTT subtitle format (amazon.transcribe; not amazon.transcribe-medical nor Bedrock models); rejected with stream=true
Language
Language specification ISO-639-1 language codes
languages (expected languages) Expected-language list (ISO-639-1) for multi-language audio; cannot be combined with language. Drives Amazon Transcribe multi-language identification; folded into the transcription context on Bedrock models
Auto language detection Automatic identification
Detected languages in response The json response reports the detected language(s) as a languages array (Amazon Transcribe only)
Streaming
stream (SSE streaming) Set stream: true to receive incremental results as server-sent events. srt and vtt are rejected rather than answered without their cues; verbose_json is accepted but degrades to text-only events — timestamps and segments are dropped
transcript.text.segment events Speaker segments streamed as they are finalized, with response_format=diarized_json on a model that labels speakers
Advanced
timestamp_granularities Word or segment level; requires response_format=verbose_json (Amazon Transcribe only)
Speaker diarization Automatic speaker separation; requires response_format=diarized_json (Amazon Transcribe only)
known_speaker_names Accepted but ignored — diarization falls back to generic speaker labels
known_speaker_references Accepted but ignored — diarization falls back to generic speaker labels
chunking_strategy Only auto is accepted; other values are rejected
temperature Bedrock models only; rejected by Amazon Transcribe
prompt Bedrock models only; rejected by Amazon Transcribe
keywords Bedrock models only (folded into the transcription context); rejected by Amazon Transcribe — use a pre-created custom vocabulary via the VocabularyName extra parameter
include (logprobs) Accepted on Bedrock models but never populated (logprobs is always null); rejected by Amazon Transcribe
Extra model-specific params Amazon Transcribe optional settings via JSON body (see below)
Usage tracking
Input audio duration Seconds (billing unit on Amazon Transcribe)
Output text tokens On models from Bedrock

Legend:

  • Supported — Fully compatible with OpenAI API
  • Available on Select Models — Check your model's capabilities
  • Partial — Supported with limitations
  • Unsupported — Not available in this implementation
  • Extra Feature — Enhanced capability beyond OpenAI API

Models

Amazon Transcribe Amazon Models

Model Supported Languages Notes
amazon.transcribe 100+ Full-featured transcription with speaker diarization and subtitle generation at the cost of higher latency
amazon.transcribe-medical US English Clinical dictation and patient–clinician conversations, with medical terminology — see Medical transcription

Configuration Required

You must configure a bucket to use these models, through AWS_S3_BUCKET, AWS_TRANSCRIBE_S3_BUCKET, or an AWS_S3_REGIONAL_BUCKETS entry for a region where Amazon Transcribe is a candidate. This bucket is used for temporary storage during transcription processing.

Mistral Mistral Models

Model Supported Languages Notes
mistral.voxtral-mini-3b-2507 Multilingual (auto-detected) Compact model for fast transcription
mistral.voxtral-small-24b-2507 Multilingual (auto-detected) Larger model for enhanced accuracy

Mistral Voxtral Limitations

Mistral Voxtral models have the following restrictions when running on Amazon Bedrock:

  • File size limit: ~2MB maximum input file size
  • Audio channels: Mono channel audio only (single channel)

Amazon Nova Amazon Nova Sonic

Model Supported Languages Notes
amazon.nova-2-sonic-v1:0 Multilingual (auto-detected) Low-cost real-time speech recognition

Name this model to transcribe through Amazon Nova Sonic instead of Amazon Transcribe. It is the cheapest transcription option (about $0.006 per minute of audio at current Amazon Bedrock rates) and returns punctuated text in the language that was spoken. Transcription is model selection, never automatic: requests that do not name this model are unaffected.

What this model does not provide

  • Response formats: json and text only. srt, vtt, verbose_json and diarized_json are rejected, as is timestamp_granularities — this model returns no timestamps and does not report which language it detected. Use amazon.transcribe for subtitles, timestamps, speaker diarization or detected-language reporting.
  • Audio length: up to 10 minutes per request. Longer recordings are rejected; use amazon.transcribe, which has no such limit.

Other Amazon Bedrock Models

Any Amazon Bedrock model that accepts the SPEECH input modality through the Converse API can transcribe out of the box: the gateway sends the audio together with a transcription prompt and returns the model's text output.

Audio Input Formats on Bedrock Models

Uploads in the formats the Bedrock Converse audio block accepts — aac, flac, m4a, mka, mkv, mp3, mp4, mpeg, mpga, ogg, opus, pcm, wav, webm, and x-aac — are sent through as-is. Any other audio or video upload is automatically converted to FLAC before transcription (requires FFmpeg on the server), including the audio track of a video container. An upload that is neither audio nor video is rejected with the list of accepted formats; an audio or video file whose track cannot be decoded is rejected as carrying no decodable audio.

Working with Amazon Transcribe

Amazon Transcribe Amazon Transcribe Features

Model & Features:

  • Use amazon.transcribe with the same interface as OpenAI's Whisper API
  • Or use OpenAI model names directly: whisper-1, gpt-transcribe, gpt-live-transcribe, gpt-4o-transcribe, and gpt-4o-mini-transcribe work out of the box (they map to amazon.transcribe)
  • Auto-detect language or specify it for faster processing, or list the expected languages with languages for multi-language audio
  • Word-level or segment-level timestamps with verbose_json
  • Speaker Diarization : Automatically identify and label different speakers with diarized_json, in one response or streamed segment by segment. Each speaker gets one of the sequential capital letters A, B, ... in the order the speakers are first heard — rolling over to AA, AB, ... past 26 speakers — and keeps it for the whole transcript
  • Native Subtitles : SRT/VTT files generated directly by Amazon Transcribe with precise timing

OpenAI Model Compatibility

stdapi.ai includes built-in model aliases that map the OpenAI model names to Amazon Transcribe:

  • whisper-1 → amazon.transcribe
  • gpt-transcribe → amazon.transcribe
  • gpt-live-transcribe → amazon.transcribe
  • gpt-4o-transcribe → amazon.transcribe
  • gpt-4o-mini-transcribe → amazon.transcribe

These aliases let an OpenAI-based tool reach these models under the names it already sends, with no configuration change. You can also customize or override these aliases to suit your needs.

Performance Tips: Optimize Speed & Cost

  • Specify the language if you know it—skips auto-detection for faster processing and lower AWS costs

Provider-Specific Parameters

Unlock advanced Amazon Transcribe capabilities by passing provider-specific parameters directly in your request body. These parameters are forwarded to Transcribe's StartTranscriptionJob API.

JSON body required

Unlike the multipart/form-data upload, extra parameters are only reachable through the application/json request body (file as base64, data URI, HTTPS URL, or file-id: reference) — the multipart path only accepts the documented OpenAI fields.

PII Redaction:

Redact personally identifiable information from the transcript (only the single-output redacted mode is supported):

{
  "model": "amazon.transcribe",
  "file": "data:audio/mp3;base64,<base64-encoded-audio>",
  "ContentRedaction": {
    "RedactionType": "PII",
    "PiiEntityTypes": ["NAME", "SSN", "CREDIT_DEBIT_NUMBER"]
  }
}

Custom Vocabulary and Filtering:

Improve recognition of domain-specific terms and mask or remove profanity/sensitive words:

{
  "model": "amazon.transcribe",
  "file": "data:audio/mp3;base64,<base64-encoded-audio>",
  "VocabularyName": "MyCustomVocabulary",
  "VocabularyFilterName": "MyProfanityFilter",
  "VocabularyFilterMethod": "mask"
}

Alternative Transcriptions and Channel Identification:

Request multiple candidate transcriptions per segment, or transcribe each audio channel separately (e.g. two-party phone calls recorded in stereo):

{
  "model": "amazon.transcribe",
  "file": "data:audio/mp3;base64,<base64-encoded-audio>",
  "ShowAlternatives": true,
  "MaxAlternatives": 3,
  "ChannelIdentification": true
}

Incompatible with diarized_json

ChannelIdentification cannot be combined with response_format=diarized_json, which already forces AWS speaker-label diarization. Requesting both returns HTTP 400.

Toxicity Detection:

Flag toxic content (profanity, hate speech, harassment) in the transcript:

{
  "model": "amazon.transcribe",
  "file": "data:audio/mp3;base64,<base64-encoded-audio>",
  "ToxicityDetection": [{"ToxicityCategories": ["ALL"]}]
}

Multi-Language Identification:

Detect and transcribe multiple languages spoken in the same audio, optionally restricted to a candidate list:

{
  "model": "amazon.transcribe",
  "file": "data:audio/mp3;base64,<base64-encoded-audio>",
  "IdentifyMultipleLanguages": true,
  "LanguageOptions": ["en-US", "es-US", "fr-FR"]
}

Per-Language Custom Resources:

A custom vocabulary, vocabulary filter or custom language model can only be attached per candidate language when the language is identified rather than given. Pair LanguageIdSettings with LanguageOptions so the dialect your resources were created for is the one identified:

{
  "model": "amazon.transcribe",
  "file": "data:audio/mp3;base64,<base64-encoded-audio>",
  "LanguageOptions": ["en-US", "es-US"],
  "LanguageIdSettings": {
    "en-US": {"VocabularyName": "MedicalTermsEnUs"},
    "es-US": {"VocabularyName": "MedicalTermsEsUs"}
  }
}

Identification required

LanguageIdSettings applies to identified languages only. Combined with a fixed language (or a single-entry languages), the request returns HTTP 400 rather than silently dropping the custom resources — use the flat VocabularyName, VocabularyFilterName and ModelSettings parameters in that case. LanguageModelName is not available with IdentifyMultipleLanguages.

Standard languages parameter

The standard OpenAI languages parameter drives the same multi-language identification with plain ISO-639-1 codes (e.g. ["en", "es", "fr"]), works on the multipart path too, and the detected language(s) come back in the json response's languages array. A single-entry list behaves like language. Do not combine it with language or with the provider-specific parameters above.

Configuration Options:

Option 1: Per-Request

Add provider-specific parameters directly in your JSON request body (as shown in examples above).

Option 2: Server-Wide Defaults

Configure default parameters for amazon.transcribe via the DEFAULT_MODEL_PARAMS environment variable:

export DEFAULT_MODEL_PARAMS='{
  "amazon.transcribe": {
    "VocabularyFilterName": "MyProfanityFilter",
    "VocabularyFilterMethod": "mask"
  }
}'

Note: Per-request parameters override server-wide defaults.

Behavior:

Compatible parameters are forwarded to Amazon Transcribe and applied; unsupported parameters or values return HTTP 400 with an error message.

Available Parameters:

The following parameters from Amazon Transcribe's StartTranscriptionJob API can be used:

  • ContentRedaction (object): PII redaction — RedactionType (PII), PiiEntityTypes (list), RedactionOutput (redacted only; redacted_and_unredacted is rejected — the unredacted copy is not tracked for automatic cleanup)
  • VocabularyName (string): Custom vocabulary to improve recognition accuracy
  • VocabularyFilterName / VocabularyFilterMethod (string / mask, remove, tag): Profanity or sensitive-word filtering
  • ShowAlternatives / MaxAlternatives (bool / integer 2-10): Return multiple candidate transcriptions per segment
  • ChannelIdentification (bool): Transcribe each audio channel separately (incompatible with diarized_json)
  • MaxSpeakerLabels (integer 2-30): Maximum speakers to identify with response_format=diarized_json (default 10). It has no effect on a streamed transcription delivered phrase by phrase (see Streaming). Amazon Transcribe distinguishes at most 30 speakers whatever the value, and separates them less reliably past about five
  • ShowSpeakerLabels (bool): Always on with response_format=diarized_json; setting it directly with another format runs AWS speaker labeling without exposing speaker data in the response
  • ToxicityDetection (list): Toxic-content flagging — [{"ToxicityCategories": ["ALL"]}]
  • IdentifyMultipleLanguages / LanguageOptions (bool / list): Multi-language identification, optionally restricted to a candidate list (supersedes language; cannot be combined with the standard languages parameter)
  • LanguageIdSettings (object): Per-language VocabularyName, VocabularyFilterName and LanguageModelName, keyed by language code (up to five) — the only way to attach them when the language is identified rather than given
  • ModelSettings (object): LanguageModelName — custom language model selection

VocabularyName, VocabularyFilterName, and custom language models must already exist in your AWS account (created via the AWS Transcribe console, CLI, or SDK) before being referenced here.

Medical transcription

amazon.transcribe-medical transcribes clinical audio — a physician's dictated notes, a patient–clinician conversation — with Amazon Transcribe Medical, which recognizes medical terms, drug names and dosages. It is priced separately from amazon.transcribe (see Amazon Transcribe pricing) and needs its own IAM permissions.

{
  "model": "amazon.transcribe-medical",
  "file": "data:audio/wav;base64,<base64-encoded-audio>",
  "Type": "DICTATION"
}
Parameter Values Notes
Type CONVERSATION (default), DICTATION One speaker dictating, or a conversation between several
Specialty PRIMARYCARE (default), CARDIOLOGY, NEUROLOGY, ONCOLOGY, RADIOLOGY, UROLOGY Any value other than PRIMARYCARE requires stream=true (without it, HTTP 400) and a deployment serving live medical transcription (without it, HTTP 503)
VocabularyName string A medical custom vocabulary you created
ShowSpeakerLabels, MaxSpeakerLabels, ChannelIdentification, ShowAlternatives, MaxAlternatives as for amazon.transcribe Not carried by a phrase-by-phrase stream, so combining them with Specialty returns HTTP 400
  • Language: US English only. language and languages may be omitted or set to en; any other language returns HTTP 400, as does any amazon.transcribe parameter this model does not list above (redaction, toxicity detection, language identification, vocabulary filters, custom language models, ContentIdentificationType).
  • Formats: json, text, verbose_json and diarized_json. srt and vtt return HTTP 400: request verbose_json for timed segments.
  • Streaming: stream=true streams phrase by phrase, since the language is always known. A request setting a parameter a stream cannot carry, or reaching a deployment where no configured region offers live medical transcription, is still served, all at once when the transcript is complete — except with a Specialty other than PRIMARYCARE, which needs the live path (HTTP 503 without it).
  • Translation: not available — the transcript is already in English.
  • Regions: medical transcription is offered in fewer regions than amazon.transcribe, and fewer still when streamed (supported regions). A request moves on to the next configured region that offers it; when none does, it returns HTTP 503.
  • Fixed variants: a client that can only send a model name (Home Assistant, Open WebUI) reaches a specific audio Type through an alias that carries configuration. A Specialty other than PRIMARYCARE in an alias only works for clients that send stream=true; the others get HTTP 400.

    export MODEL_ALIASES='{
      "medical-dictation": {
        "model": "amazon.transcribe-medical",
        "extra_params": {"Type": "DICTATION"}
      }
    }'
    

Protected health information

Amazon Transcribe Medical is a HIPAA-eligible service; processing protected health information requires a Business Associate Addendum with AWS and a deployment configured for it. A request served as a transcription job — every non-streamed request, and a streamed one that falls back to a job — writes the audio and the transcript under the AWS_S3_TMP_PREFIX of your transcription bucket and deletes them after responding, best effort; a lifecycle rule expiring that prefix removes whatever an interrupted server leaves behind. In a versioned bucket a deletion keeps a noncurrent version: the buckets the Terraform module creates expire both current and noncurrent versions under that prefix after one day, so the data can remain for a day or two; a bucket you bring needs the same noncurrent-version expiration on that prefix, or deleted audio and transcripts stay recoverable indefinitely. Transcripts also reach the server log when LOG_REQUEST_PARAMS is enabled — keep it off for this model.

Streaming

stream=true returns the transcript as server-sent events — transcript.text.delta events followed by a final transcript.text.done — instead of one response body. A ready-to-run example is below.

Each phrase is sent as it is recognized, rather than after the whole recording, whenever the request names the language to expect: send language, or two or more expected languages. That path needs no S3 bucket, so a deployment with no storage configured serves streamed transcriptions.

When no configured region can open a live session — the deployment lacks the transcribe:StartStreamTranscription permission (transcribe:StartMedicalStreamTranscription for the medical model), or no region offers streaming for the model — the request is served as a transcription job instead: it is still streamed, but its events arrive together at the end, and it needs a transcription bucket.

A request naming neither is still streamed, but its events arrive together once the recording has been read and its language detected. Operators can set AWS_TRANSCRIBE_STREAM_LANGUAGES to the languages their callers actually send, which gives those requests the faster path too. The same applies to a request using any provider-specific parameter above other than VocabularyName, VocabularyFilterName and VocabularyFilterMethod, which are the only ones a phrase-by-phrase transcript can carry.

Speaker segments

With response_format=diarized_json, each stretch of transcript spoken by one speaker is also sent as a transcript.text.segment event, carrying the segment's id, start, end, speaker and text. Segments are interleaved with the deltas and always precede the final transcript.text.done.

A speaker is attached only once the recognizer has settled the words it labels, so a segment is sent once and its speaker never changes afterwards. Segments carry the same speaker labels as the non-streamed diarized_json response. Naming the language is what gets the segments phrase by phrase instead of together at the end: send language, or two or more expected languages.

Each delta belonging to a segment names that segment's id in its segment_id field, so the two can be correlated without matching their text. Text no speaker was attributed to — including text the recognizer never settled — arrives as a delta with no segment_id and no segment event of its own, so the concatenated deltas — not the segments — are the complete transcript.

A model that reports no speakers rejects diarized_json with stream=true rather than answering with an unlabelled transcript.

Streamed events carry no subtitle cues, which is why srt and vtt are rejected with stream=true.

Limits and behaviour to know

  • With amazon.transcribe, the prompt, temperature, keywords and include parameters are rejected with an error to ensure consistent transcription accuracy. For keywords, the error points at the pre-created custom vocabulary alternative via the VocabularyName extra parameter.
  • known_speaker_names and known_speaker_references are accepted but ignored for every model: Amazon Transcribe's automatic speaker diarization runs without known speaker references, so it falls back to generic speaker labels.
  • include: ["logprobs"] is accepted on Bedrock models but never populated — logprobs comes back null — because the Converse API returns no token log probabilities.
  • With amazon.transcribe, the decoder fields of a verbose_json segment are placeholders: seek is always 0, temperature always 0.0 and tokens always empty, while avg_logprob and no_speech_prob report the documented silence signal (-2.0 and 1.0 on a segment carrying no speech, 0.0 and 0.0 otherwise) rather than measured probabilities. The model reports no per-segment decoder state, so start, end and text are the values that carry information: a transcript cannot be re-aligned or re-decoded from the rest.
  • chunking_strategy accepts auto only; any other value is rejected rather than silently applied.
  • amazon.transcribe-medical accepts US English only, refuses srt and vtt, and takes a specialty other than primary care only with stream=true on a deployment serving live medical transcription — see Medical transcription.
  • The extra Amazon Transcribe parameters are reachable through the application/json body only, not through the multipart upload — see Provider-Specific Parameters.

Request headers

This endpoint supports standard Bedrock headers for enhanced control over your requests. All headers are optional and can be combined as needed.

Content Safety (Guardrails)

Header Purpose Valid Values
X-Amzn-Bedrock-GuardrailIdentifier Guardrail ID for content filtering Your guardrail identifier
X-Amzn-Bedrock-GuardrailVersion Guardrail version Version number (e.g., 1)

The guardrail evaluates the transcript the model produced, not the audio sent. On a streamed request the events are withheld until the transcript is complete, so the guardrail sees the whole text before any of it is delivered. When it masks content, the stream carries the masked transcript as a single delta and the speaker segments are withheld, each one repeating text the guardrail took out — so a streamed diarized_json request answers with the masked transcript, where the same request without stream=true fails with content_filter instead. X-Amzn-Bedrock-Trace is accepted but has no effect on this route — no guardrail trace is returned.

Performance Optimization

Header Purpose Valid Values
X-Amzn-Bedrock-Service-Tier Service tier selection default, flex, priority, reserved
X-Amzn-Bedrock-PerformanceConfig-Latency Latency optimization standard, optimized

Both apply only to models transcribed through the Amazon Bedrock runtime. They have no effect on amazon.transcribe, which is served by Amazon Transcribe, nor on Amazon Nova Sonic, which is served over its own bidirectional stream.

Example with headers:

curl -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "X-Amzn-Bedrock-GuardrailIdentifier: your-guardrail-id" \
  -H "X-Amzn-Bedrock-GuardrailVersion: 1" \
  -F file=@meeting-recording.mp3 \
  -F model=amazon.transcribe \
  -F response_format=json

Detailed Documentation

For complete information about these headers, configuration options, and use cases, see:

Try it

Transcribe audio to JSON:

curl -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F file=@meeting-recording.mp3 \
  -F model=amazon.transcribe \
  -F response_format=json

Transcribe via JSON body (MCP and AI agents):

When using MCP tools or HTTP clients that cannot construct multipart requests, pass the audio as a data URI or URL:

# Data URI (inline base64)
curl -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "file": "data:audio/mp3;base64,<base64-encoded-audio>",
    "model": "amazon.transcribe",
    "response_format": "json"
  }'
# HTTPS URL (server fetches the audio)
curl -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "file": "https://example.com/audio.mp3",
    "model": "amazon.transcribe"
  }'
# Files API reference (file-id: URI scheme)
curl -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "file": "file-id:file-0190c51c7de7455d9b8c2efe27dfbf67",
    "model": "amazon.transcribe"
  }'

See Files API → Referencing Uploaded Files for the full description of the file-id: URI scheme.

Generate subtitles:

curl -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F file=@video-audio.mp3 \
  -F model=amazon.transcribe \
  -F response_format=srt \
  -F language=en

Transcribe with speaker diarization:

curl -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F file=@meeting-recording.mp3 \
  -F model=amazon.transcribe \
  -F response_format=diarized_json

Stream a transcription as SSE events:

Set stream=true to receive the transcript incrementally as server-sent events (transcript.text.delta events followed by a final transcript.text.done event):

curl -N -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F file=@meeting-recording.mp3 \
  -F model=amazon.transcribe \
  -F language=en \
  -F stream=true

Stream speaker segments as they are recognized:

curl -N -X POST "$BASE/v1/audio/transcriptions" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F file=@meeting-recording.mp3 \
  -F model=amazon.transcribe \
  -F response_format=diarized_json \
  -F language=en \
  -F stream=true

Each completed speaker segment arrives as a transcript.text.segment event beside the text deltas:

data: {"delta":"Good morning, how can I help?","type":"transcript.text.delta","segment_id":"seg_0"}

data: {"id":"seg_0","start":0.01,"end":4.27,"speaker":"A","text":"Good morning, how can I help?","type":"transcript.text.segment"}

Naming the language starts the transcript sooner

language is what lets each phrase be sent as it is recognized rather than after the whole recording — see Streaming Transcriptions.

verbose_json streams as plain text

stream=true combined with response_format=verbose_json is accepted rather than rejected, but the streamed events carry transcript.text.delta / .done only — segment timings, word timings and language details are not included. Request verbose_json without stream to get them.

Next steps

Next: Models API · Speech to English API · AWS_TRANSCRIBE_S3_BUCKET, where non-streamed jobs are staged · Speech-to-text IAM permissions