Releases & Roadmap¶
stdapi.ai is under active development with regular feature releases.
Version Index¶
Latest: v1.19.1 — released 2026-09-29.
Every release, newest first. Each entry in the Release History below opens with a five-bullet summary.
| Version | Date | Theme | Release notes |
|---|---|---|---|
| v1.19.0 (and v1.19.1) | 2026-09-23 (2026-09-29) | Long conversations, token counting on every model, reasoning control, medical transcription | Read |
| v1.18.0 | 2026-09-19 | Vendor parity across every mirrored API, tenant key rotation, per-tenant rate limits, realtime tools | Read |
| v1.17.0 | 2026-09-08 | Your own model endpoints, the Ollama dialect, per-tenant keys, usage and cost from the API, WebRTC | Read |
| v1.16.0 (and v1.16.1) | 2026-08-21 (2026-08-25) | Conversations, batches, vector stores, realtime speech and per-user identity | Read |
| v1.15.0 | 2026-08-03 | Reliability, performance and feature completeness | Read |
| v1.14.0 | 2026-07-12 | Bedrock Mantle, video generation, Cohere APIs, moderation and stored conversations | Read |
| v1.13.0 | 2026-07-03 | Terraform module compliance and security hardening | Read |
| v1.12.0 | 2026-05-29 | Completions API, video understanding and file references | Read |
| v1.11.0 (through v1.11.4) | 2026-05-02 (2026-05-28) | MCP server, agent discovery and model search | Read |
| v1.10.0 | 2026-04-17 | OpenAI Responses API | Read |
| v1.9.0 | 2026-04-10 | Files API and a JSON body for the Images API | Read |
| v1.8.0 | 2026-04-04 | Broader model compatibility and structured output | Read |
| v1.7.0 | 2026-03-20 | Automatic region routing, deprecated model fallback and resilience | Read |
| v1.6.0 | 2026-02-27 | Anthropic API compatibility and advanced Claude capabilities | Read |
| v1.5.0 (and v1.5.1–v1.5.2) | 2026-02-15 (2026-02-18) | Advanced reasoning and model compatibility | Read |
| v1.4.0 | 2026-02-11 | Audio enhancements and model compatibility | Read |
| v1.3.0 (through v1.3.5) | 2026-01-11 (2026-02-02) | Image editing and variation support | Read |
| v1.2.0 | 2025-12-18 | Service tiers, system tools and performance | Read |
| v1.1.0 | 2025-11-27 | Embeddings, prompt caching and advanced routing | Read |
| v1.0.0 | 2025-11-10 | Foundation release | Read |
Roadmap (Tracked on GitHub)¶
Pending features and current deployment state are tracked on the GitHub Project.
Release History¶
v1.19.0 – 2026-09-23 – Long Conversations, Reasoning Control & Medical Transcription (with v1.19.1 maintenance update, 2026-09-29)¶
At a glance
- Long conversations keep going. The Responses API serves
truncation: "auto",context_managementcompaction and thecompaction_triggeritem, where all three answered400or were dropped. On Anthropic Messages, Claude clears old tool results and thinking blocks throughcontext_managementand reports what it cleared. - A full context window is refused the way the vendor refuses it.
400context_length_exceededon the OpenAI routes,prompt is too longon Anthropic Messages, before the first byte of a stream, and never in the backend's own words. A malformed request is refused the way the same vendor endpoint refuses it, with the parameter and code where the vendor names them. - Token counting answers for every text model.
input_tokensandcount_tokenscount past the context window, with server tools, images and documents, on every model including those served by Bedrock Mantle and your own endpoints. Claude up to 4.6 is counted exactly; the others get an approximation that errs high, never low. - Reasoning does what the request asks. The effort level now reaches OpenAI GPT-5.x, GPT-6, gpt-oss and Moonshot Kimi K3, which were dropping it; recent Claude models return their reasoning when asked for a summary; and Claude Code's interleaved thinking is no longer filtered out.
- Medical transcription.
amazon.transcribe-medicaltranscribes clinical dictation and patient–clinician conversations in US English, streamed phrase by phrase, with six medical specialties.
This release is about what happens when a conversation gets long. An agent session eventually outgrows its model's context window, and the gateway had one answer for that: a refusal, often in the backend's words rather than the API's. The Responses API now trims or compacts a conversation when asked, Claude clears stale tool results through Anthropic's context editing, an overflow that still happens is refused exactly as the vendor refuses it, and a client can count tokens on every text model to see it coming. Around that: reasoning controls that now reach the model families that were dropping them, GPT-6 web search and code interpreter with no configuration, Moonshot Kimi K3, medical transcription, and a pass over reported costs that corrects several prices, one of them reported up to 50% above what AWS bills.
New IAM Permissions, All Optional
Nothing new is required. A deployment that does not serve medical transcription needs the policy it already has; the Terraform module grants these wherever it grants transcription. See Speech-to-Text.
- Medical transcription:
transcribe:StartMedicalTranscriptionJobandtranscribe:StartMedicalStreamTranscriptionon*(neither accepts a resource type),transcribe:GetMedicalTranscriptionJobandtranscribe:DeleteMedicalTranscriptionJobonmedical-transcription-job/*, andtranscribe:TagResourceextended to that same resource. Without themamazon.transcribe-medicalis unavailable, and standard transcription is unaffected.
Behavior Changes
Review these before upgrading. They may change what existing clients observe:
- The OpenAI GPT-6 family is served by Bedrock Mantle by default, wherever a configured Mantle region lists it, so its
web_searchandcode_interpreterwork with no configuration.AWS_BEDROCK_MANTLE_PREFERRED_MODELSnow defaults toopenai.gpt-5.6,openai.gpt-6. It is the trade GPT-5.6 already made: no Global routing discount, so GPT-6 Astra costs exactly 10% more per token, no guardrails, and usage billed and reported under Bedrock Mantle. An empty value keeps both families on the classic endpoint. An explicit value replaces the default, so a deployment that already sets one keeps GPT-6 where it was. - A context-window overflow is refused in the API's own terms. Chat Completions answers
400context_length_exceedednamingmessages, Responses the same naminginput, Anthropic Messagesinvalid_request_errorprompt is too long: N tokens > M maximum, and legacy Completions and Ollama their own wording, on every model, Bedrock Mantle included. A streamed Chat Completions, Completions, Messages or Ollama request that the model refuses on its first event now gets that400before any byte, instead of a200followed by an error event. A streamed Responses request keeps the in-streamerrorthenresponse.failed, as OpenAI streams it, and thaterrorevent now also nests the error undererror, as the live API does, so the official SDK raises it. - Validation errors are worded as each vendor words them. On the OpenAI routes
error.paramnames the field in OpenAI's notation (messages[0].content) anderror.codethe failure (unknown_parameter,missing_required_parameter,invalid_type,invalid_valueand the range codes); Anthropic routes answer<field>: <reason>with no gateway prefix. A client matching on the old message text needs updating. See Using stdapi.ai. STRICT_INPUT_VALIDATIONrefuses an unknown field only where OpenAI does. It is now ignored on chat messages other thandeveloper, on content parts, tool calls and tool definitions, on Responses function tools and on moderation requests. A replayed tool call carrying an unknown field is now ignored whatever the setting.- Token counting no longer answers
400for any text model. Models served by Bedrock Mantle (the GPT-5.6 and GPT-6 families by default), Marketplace and SageMaker AI endpoints, inputs past the context window and requests offering server tools are all counted, andsearch_modelslists both counting routes for every text model. Most of these counts are approximate: see Known Limitations below. reasoning.summaryon the Responses API now returns a summary. Any value puts the reasoning in the item'ssummaryassummary_textparts, streamed asresponse.reasoning_summary_*events, where it was ignored and the text left incontent. Codex, which displayssummaryalone by default, now shows the reasoning.- A requested reasoning effort now takes effect on OpenAI GPT-5.x, GPT-6 and gpt-oss, where it was silently dropped, so their output-token usage follows the level you ask for rather than the model's default.
minimal, which none of them accepts, is sent aslow. - Transcription usage is recorded to the second, with no 15-second minimum, matching AWS's published billing: a 2-second clip was recorded as 15. A streamed transcription is costed at the streaming rate and carries
input_seconds_by_spec: {"streaming": n}, where the batch rate under-reported it by 40% inus-east-1. - A streamed transcription no live session can serve falls back to a transcription job rather than answering
503. When no configured region opens a live session, for lack of thetranscribe:StartStreamTranscriptionpermission or of a region that streams the model, the request is still streamed, but its events arrive together at the end and it needs a transcription bucket. - A bare
computertool on Claude Opus 5 and 5.5 is now the computer-use server tool, as on the other current Claude models, where it was passed through as a custom tool. - Kimi K3 is no longer listed as Batch API capable: AWS refuses batch inference for it.
- Compaction item IDs now start with
cmp_, as upstream's do.
Known Limitations
- Most token counts are approximate, and err high. A count is exact for Claude up to 4.6 and for the Claude models served by Bedrock Mantle. Claude 4.7 and later (Opus 4.7, Opus 5.5, Sonnet 5, Fable) are counted 7% to 70% above the exact figure, about 50% typically. Every other model, the Mantle-served GPT families and your own endpoints included, is counted at least 5% above it and about 30% typically, more for one-line prompts and non-Latin scripts. Across 27 tokenizers and everything measured, no approximation came in below the exact count. On those models
compact_thresholdis approximate too, andtruncation: "auto"counts the whole input. A file given by URL, S3 URI or file ID is counted from its size and never downloaded. The usage a response reports is always exact. See Input Token Counting. - An exact count can drift slightly past the context window and on a history holding web search results: within 20 tokens on a 300,000-token message, and within 2% with search results, on Claude Haiku 4.5.
- Recent Claude models return no reasoning text on Chat Completions. Opus 4.7 and later, Sonnet 5, Fable and Mythos return it only as a summary, which Chat Completions cannot request: use
thinking.displayon Messages orreasoning.summaryon Responses. - Context editing is Claude-only, and has no compaction. Other models accept
context_managementand ignore it with a server-log warning,compact_20260112is refused as an unknown edit type, and batch results do not reportapplied_edits. - Compaction items are not encrypted. They carry the conversation's text, system and developer messages included: hand them only to parties that may read and alter that conversation, or continue with
previous_response_idto keep it on the server. - GPT-6 Sol and Luna have no published AWS price yet, so their usage is recorded without a cost. Mantle lists GPT-6 in few regions: where no configured Mantle region does,
web_searchandcode_interpreterare refused with400. - Medical transcription is US English only, refuses
srtandvtt(verbose_jsoncarries timed segments), and takes a specialty other than primary care only when streamed on a deployment serving live medical transcription.
New API Features¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
truncation: "auto" on Responses – an input over the context window has its oldest turns dropped, then the latest turn's longest text cut at its end, until it fits. System and developer messages are always kept, usage reports the input that was kept, and input_tokens counts what a response would keep | ||
context_management compaction on Responses – once the input crosses compact_threshold, everything but the latest step is summarized into a compaction item that leads output and stands for the whole history when sent back; streamed at index 0 before the answer. A trailing compaction_trigger item compacts on demand, where it was silently dropped. Both calls are billed, and usage adds them up | ||
The output is capped when only it overflows – an input the window holds alone but not beside max_output_tokens or max_tokens is answered with a shorter output on Responses and Anthropic Messages, as both vendors answer it. Chat Completions keeps refusing it, as OpenAI does | ||
Token counting on every text model – /v1/responses/input_tokens and /v1/messages/count_tokens count an input past the context window, always above the window so a client learns the conversation no longer fits, plus server tools with their replayed calls, results and citations, inline images, and documents and remote files at an allowance | ||
Context editing – clear_tool_uses_20250919 and clear_thinking_20251015 with every documented option. Claude applies the edits and reports them in applied_edits, on the final message_delta when streaming, and count_tokens counts the edited prompt beside original_input_tokens. The beta flag is added for you, so the header is optional | ||
thinking.display – summarized or omitted reaches Claude as sent, so Opus 4.7 and later, Sonnet 5, Fable and Mythos, which omit their thinking text by default, return a summary when asked. The thinking tokens bill the same either way. On Responses, reasoning.summary asks for the same summary | ||
Medical transcription – amazon.transcribe-medical on /v1/audio/transcriptions, for a physician's dictation or a patient–clinician conversation (Type), with Specialty and medical custom vocabularies. json, text, verbose_json and diarized_json are served, and stream=true streams phrase by phrase. A client that can only send a model name reaches a fixed variant through a model alias |
Models & Reasoning¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| GPT-6 web search and code interpreter – served with no header and no configuration wherever a configured Mantle region lists the model, and Chat Completions reaches GPT-6 natively there rather than converted | ||
Reasoning effort on GPT-5.x, GPT-6 and gpt-oss – from every dialect: reasoning_effort, reasoning.effort, output_config.effort or thinking. none turns reasoning off where the model can; GPT-6 Astra, and gpt-oss outside Bedrock Mantle, cannot, and are served at their default level with a warning in the request log. gpt-oss gets its own low / medium / high scale | ||
Kimi K3 – reasoning effort from every dialect, with none answering without reasoning, and automatic prompt caching: a repeated prefix is reported as cached tokens on Chat Completions, Responses and Messages and costed at the cache rates. Cache breakpoints are accepted and ignored, since the model rejects them. A turn replaying the model's own reasoning is served | ||
Claude Opus 5.5 – a request disabling reasoning on Opus 5.5, which always reasons, is served at its adaptive default with a warning in the request log, as on Fable and Mythos, rather than refused. On Bedrock Mantle, xhigh and max reach Claude 4.7 and later unchanged instead of being capped to high |
Cost Reporting¶
- Daybreak Blue (GPT-5.6 Sol) was reported above what AWS bills: $5.50 / $33.00 per million input / output tokens against AWS's $4.40 / $22.00, and $11.00 / $49.50 against $8.80 / $33.00 past 272K, with the cache rates 25% high as well. It is now costed at GPT-5.6 Sol's rates.
- GPT-5.4 and GPT-5.5 are costed at their long-context rates past 272K input tokens, the boundary AWS publishes for them, where every call was costed at the short rate. See Long-Context Pricing.
- GPT-6 Astra is priced: In-Region $11.00 / $55.00 and Global $10.00 / $50.00 per million input / output tokens, with cache and long-context rates. It had no price at all.
- Kimi K3 calls are priced, where every call reported no cost. A Global call is costed at its Global rate ($3.00 / $15.00) rather than the US one ($3.30 / $16.50), and cache reads and writes are priced. A Grok 4.6 Global call now reaches its Global rate ($2.00 / $6.00) the same way.
Platform Features¶
- Transcription moves on from a region that does not offer the operation. Medical transcription and live streaming run in fewer regions than standard transcription; a region lacking one now hands the request to the next configured region instead of failing it, and when none offers it the answer is a
503. See Resilience. - The Models page lists Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna (the last two without a price, as AWS publishes none yet), GPT-6 Astra with its prices, the GPT-5.4 and GPT-5.5 long-context tier, and the Global rates of Kimi K3 and Grok 4.6.
Fixes¶
- Claude Code's interleaved thinking reaches Claude. The allowlist carried the interleaved-thinking flag capitalised, which no client sends, so thinking between tool calls silently stayed off. Separately, a request carrying the memory tool or a computer-use tool lost every flag of its
anthropic-betaheader, the 1M context window's included, when the gateway added the one that tool needs; the two lists are now joined, andcount_tokensmerges and filters the flags the same way. - Reasoning effort on Claude 3.7 to 4.5 served by Bedrock Mantle no longer fails: it was sent as an effort level these models refuse, so every such request answered
400. It now becomes the thinking budget they take, scaled on the output limit; an output limit of 1,024 tokens or less leaves no room for one, and the request is served without reasoning and a logged warning. - A context-window overflow no longer leaks the backend's error text, which on some models named an internal service and a request identifier, and a streamed one is no longer reported as a retryable
server_error. - A failed Responses request on a Bedrock Mantle model converted to that API is refused, where it was answered as an empty success.
Fixes & Maintenance (v1.19.1)¶
- Claude Sonnet 5.5 is served without errors. Turning reasoning off or forcing a tool surfaced a raw
400from Amazon Bedrock. Thinking turned off on any route (reasoning_effort: "none",enable_thinking: false,reasoning.effort: "none",thinking: {"type": "disabled"},think: false) now reaches it asbetween_tools, its lowest setting, and a forced tool choice is refused with a message namingauto, as on Opus 5.5. See Turning Thinking Off on Claude Sonnet 5.5. - Anthropic Messages accepts
thinking: {"type": "between_tools"}, as Anthropic does, on message creation and token counting. thinking: {"type": "disabled"}beside anoutput_config.effortturns thinking off and keeps the effort, as upstream does. Claude reasoned at that effort instead, and was billed for it.- Grok 4.7 calls are priced, at Global $2.00 / $6.00 per million input / output tokens. Every call was reported as free.
- GPT-6 Astra short prompts are costed at the short-context rate. The Price List rows AWS now publishes for it were read as one tier, so a short prompt could be costed at the long-context rate, up to twice as high.
- GPT-6 Luna, GPT-6 Sol and GLM 4.6 are priced, at the rates on their AWS model cards (GPT-6) and on Z.ai's pricing page (GLM 4.6). Every call was reported as free.
- The Terraform module grants the server the KMS permission its DynamoDB table needs. Since module 1.17.0, a deployment where the module created the table refused every table read and write, so tenant API keys, per-tenant rate limits and the shared model cache were unavailable. Update the module to apply it.
- A conversation replaying a turn that ended after reasoning is answered, where automatic prompt caching made Amazon Bedrock refuse it with a
400; Codex on Nova 2 Lite could not recover from a cut-short answer. - The Models page lists Claude Sonnet 5.5 and Grok 4.7, names each model once where it is served by two AWS services, quotes the nearest region's rate where the selected region has none, lists Amazon Transcribe in every region that offers it, shows closed models' licence as proprietary, and takes each model's APIs and prompt caching from its AWS model card.
v1.18.0 – 2026-09-19 – Vendor Parity, Tenant Key Rotation & Rate Limits¶
At a glance
- Requests the real APIs accept are no longer refused — about thirty of them, across Chat Completions, Responses, Anthropic Messages, the Ollama dialect, audio and images: a
temperatureabove 1,service_tier: "fast",max_tokens: 0,speedup to 4.0, an empty Ollamaformat, atool_resultwith no content, and the rest. The gateway's job is to answer what the vendor answers. - Tenant keys rotate themselves — set
TENANT_KEY_SECRETSMANAGER_PREFIXand each key becomes a Secrets Manager secret the tenant reads for itself, rotated on a schedule with an overlap window so nothing is locked out mid-flight. - Per-tenant rate limits —
requests_per_minuteandtokens_per_minuteper key or per deployment, counted across every instance. A deployment that declares none pays neither the latency nor the DynamoDB bill. - Voice agents can act — a Realtime session declares tools and answers the calls the model makes, and Chat Completions requests can ask for a server-side web search with OpenAI's own
web_search_options. - Uploads stream — an upload part is handed straight to S3 instead of being held in memory, so the per-part ceiling rises from 64 MiB to S3's own 5 GiB and a whole upload from 8 GiB to 48.8 TiB.
This release is about being the API it claims to mirror. The bulk of it is a compatibility sweep run against the vendors' own endpoints and official clients, not just against the gateway: every divergence it found is either fixed here or documented as a limitation. Alongside it, three things an operator asked for — tenant keys that rotate without anyone moving a secret by hand, per-tenant request and token limits that hold across instances, and a guardrail cost lever that scopes evaluation to the turns you name. Two capabilities round it out: tools on the Realtime API, which is what a voice agent needs to do anything beyond talking, and web search on Chat Completions, the dialect most clients actually speak.
New IAM Permissions, All Optional
Nothing new is required. A deployment that enables neither feature below needs the policy it already has; the Terraform module grants each set when its feature is turned on. See IAM Permissions.
- Tenant key rotation — five Secrets Manager actions scoped to
<prefix>/*, pluskms:GenerateDataKeyandkms:Decryptconditioned on the service and the encryption context, forTENANT_KEY_SECRETSMANAGER_PREFIX. They replace thessm:*delivery actions rather than adding to them. See Tenant Key Rotation. - Tenant rate limits —
dynamodb:UpdateItemon the shared table, in its own statement and conditioned on the counter items, so the one write-in-place action cannot touch a tenant record. Granted only when a limit is declared. See Shared Table.
Behavior Changes
Review these before upgrading — they may change what existing clients observe:
- An upload part is no longer capped at 64 MiB, and an upload is no longer capped at 8 GiB. Parts are streamed to S3 rather than buffered, so S3's limits are the only ones left: 5 GiB per part, 10,000 parts, 48.8 TiB per object. This leaves the gateway more permissive than OpenAI, which caps every part at 64 MiB — a client written to the vendor's limit is unaffected. A part sent inline as JSON keeps the 64 MiB bound, because that form is decoded whole and the number is a memory bound there. Large parts cost ephemeral disk rather than memory: the request body is spooled before it is streamed out.
- A
phaseon a non-assistant Responses input message is now refused. The official client typesphaseunder every role; the API keeps it underassistantalone and answers400unknown_parameterfor the others. The gateway accepted it silently and now answers what the vendor answers. Anullphaseis not a refusal. tool_choice: "none"is now documented as Partial, not Supported, on all three compatibility tables. The tools a turn declares are the only ones the model may call — that part was already true and is now proven by tests. What the backend cannot express is a conversation that has already called a tool: declaring no tool, ortool_choice: "none", leaves that conversation's tools callable. Nothing changed in the code; the tables stopped claiming otherwise.- An Anthropic file listing now always carries
next_page,nullon the last page, and the envelope no longer drops its null fields. This is whatanthropic.pagination.SyncPageCursorreads; before it, a client paging with the SDK saw one page and stopped. Theafter_idandbefore_idcursors still work and still page in both directions. - Three Bedrock responses now report what actually happened rather than what was asked for: the OpenAI surfaces report the service tier that served the request, a streamed Chat Completions response carries the same
moderationresults the non-streamed one returns, and a Responses object carries thetruncationthe request asked for. - The validation-error log line no longer contains the request body. A
422is logged as the field paths that failed validation, without their values, because a body can carry a credential. The line stays debuggable; there is deliberately no setting to restore full bodies. - A queued vector-store indexing job is refused for a tenant-credential key, rather than running that tenant's embeddings on the deployment's AWS account. This matches the refusal batch jobs already carry, and is a documented limitation.
- A knowledge-base vector store served search-only refuses corpus deletion. Any caller could previously delete documents from the underlying knowledge base through a store that was meant to read it.
- An expired Files API object is refused when referenced by
file_id. Expiry was honoured on every other path; afile_idinput bypassed it and kept feeding inference until the storage lifecycle rule swept the bytes. - A file of a type a vector store cannot index is refused by the attach, with a
400naming what that store does index, where it was previously accepted and settledfailedwithunsupported_filea moment later. Image, audio and video types are now unindexable on their content type alone, as the archive and office types already were, which is what the official API answers. A file attached in a batch is still reported on the file, as a batch reports everything that fails in it, and a file whose bytes turn out not to be what its content type claims still settlesfailed.
New API Features¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| Realtime function calling – a session declares tools and answers the calls the model makes, so a voice agent can do more than talk | ||
web_search_options on Chat Completions – OpenAI's only spelling for server-side web search on the dialect most clients speak, resolved through the same declaration the Responses and Messages dialects already use. A model that runs no web search refuses rather than answering from its own knowledge, which nothing in the response would distinguish | ||
Prompt-attack and personal-data moderation – jailbreak and prompt-injection detection, and personal data detection, on /v1/moderations and the moderation parameter of the chat routes, with no guardrail resource to configure | ||
echo on /v1/completions – the prompt is echoed back ahead of the completion, which the gateway can do locally and was silently dropping | ||
Opaque page tokens on GET /v1/files – next_page is served and page accepted, the way the SDK now pages. The token is opaque and refused when it is not one this API issued; the ID cursors stay | ||
expires_in_seconds on a file upload – the full upstream range (1 hour to 90 days) is honoured and expires_at reported on all three routes. Expiry is enforced on every read, so a TTL longer than the storage sweep still expires at the instant promised | ||
prompt_eval_cached_count, and the base URL's connection test – the cached-prompt count Bedrock already reports is passed through, and the base URL answers Ollama is running rather than a 404. /api/blobs/{digest} joins the store verbs this deployment refuses |
Multi-Tenancy¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| stdapi.ai | Tenant key rotation – with TENANT_KEY_SECRETSMANAGER_PREFIX each key is stored as a Secrets Manager secret the tenant reads for itself, so TENANT_KEY_ROTATION_DAYS rotates it with nobody in the loop. The superseded key keeps working for TENANT_KEY_ROTATION_OVERLAP_SECONDS — seven days by default, 0 for a hard cutover — and only the immediately previous key survives | |
| stdapi.ai | Per-tenant rate limits – requests_per_minute and tokens_per_minute, per key or as a deployment default, counted in the shared table so a limit holds across every instance. Tokens are admitted on an estimate and reconciled from the billed usage, which is what keeps a Realtime or MCP session from billing to a window that has ended. A counter that cannot be written fails closed with a 503, never a 429 the client did not earn |
Cost Control¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| stdapi.ai | AWS_BEDROCK_GUARDRAIL_SCOPE_TURNS – a guarded conversation is charged for everything the client replays, because Converse evaluates the whole history on every turn. This scopes input evaluation to the trailing user turns you name: a 6,401-character history bills 7 text units per policy and adds 409 ms, against 1 unit and 194 ms for the last turn alone. Off by default and left off unless you decide otherwise — scoping genuinely lowers detection, not just the bill: a prompt attack split across two turns is blocked today and is not blocked once only the last turn is tagged. Model output is always evaluated in full |
Platform Features¶
- Uploads stream to S3. A binary part is handed from the request straight to S3 and re-read in 8 MiB chunks for the session digest, so memory is one chunk whatever the part size. The digest keeps every guarantee it had: the fold still runs under the session lock inside a shield, and the signature is still written in the same critical section as the bytes, so a digest that does not answer for exactly the parts its signature names still cannot reach a completion.
- A batch input file is streamed end to end, rather than being materialised several times over — decoded lines, parsed bodies and a joined upload blob could each hold up to 200 MiB of one request.
- The documentation pages ship a current Swagger UI and ReDoc —
swagger-ui-dist5.33.0, which lays out only the part of a large specification that is on screen, andredoc2.5.4, which carries accessibility fixes and patched bundled dependencies./redocno longer loads anything fromcdn.redoc.ly, which the documentation had always promised it did not.
Fixes¶
- A tiny multipart field name can no longer exhaust the gateway: a bracketed index in a form field name drove an unbounded allocation, so a field a few bytes long could OOM-kill the instance. The index and the nesting depth are now bounded, above what every official client actually sends.
- An application-inference-profile ARN no longer hijacks a model's routing: resolving a model from an ARN copied the catalogue entry shallowly and then wrote through the copy, so one caller's ARN changed where every later caller in that region was routed. The lookup now copies what it mutates.
- A realtime session opened with an ephemeral client secret is guarded: the deployment's guardrail was configured on every other entry point and skipped on that one, so the untrusted-browser credential the guardrail exists for was the one case it did not cover.
- Batch creation no longer blocks every other caller, and a queued vector-store indexing job no longer strands a file or orphans its vectors when the job is interrupted or the lookup fails.
- The tenant key reconciliation loop survives an unexpected error: one failure ended it silently, and with it every later mint and rotation, until the instance was restarted.
- Audio and media stop answering with the wrong thing: a truncated encode was served as a complete
200when its input failed mid-stream, a bad upload to a live transcription never got its400, and text with an&or a<produced a BedrockInvalidSsmlExceptionwheneverspeedwas not 1.0. A transcription job that is already gone is now accepted as deleted rather than re-raising. - A file URL that is fetchable is now usable: URL ingestion follows redirects and falls back to a ranged read, where before a redirect or a server refusing a
HEADprobe failed the request outright. - A stored conversation stays readable: an
additional_toolsorcompaction_triggeritem made every later read of that conversation a500, permanently. - Requests the vendors accept are accepted:
temperatureabove 1,service_tieraliases including"fast",max_tokens: 0(the documented cache pre-warm), the 2026 browser and computer toolsets, thecontainerobject form, the custom-voice object, atool_resultblock with no content,speedup to 4.0 on the Polly engines that honour it, an explicit JSONnullon ten nullable image fields,output_compression: 0, an empty Ollamaformat(which brokelangchain-ollamaon every call), and the Ollama option sentinelsseed: -1andtop_k: 0that answered500. - Responses the vendors send are sent:
event:names on streamed image frames, the citations a streamed document answer carries,backgroundandsafety_identifieron the Converse path, thecode_interpreter_callitem the model actually ran instead of an erased one, an item reference answered with the item it names in an envelope the official SDK can parse, and the filenames the vendor keeps rather than sanitising. - A realtime turn the server detects is now committed like a turn the caller ends: on
server_vad, the API's default turn mode, the gateway sent neitherinput_audio_buffer.committednor the conversation item the turn became, so a client waiting for either against a detected turn waited for ever. - The Anthropic file listing honours
ids=, where it previously ignored it and returned the wrong set of files. - Ollama chat keeps what the caller sent: images on a tool, assistant or system message were silently dropped, and two endpoints answered an accidental
404. - The Models page states what the source measured: two benchmark scores were published against the wrong model — GLM 4.7 carried GLM 4.7 Flash's GPQA Diamond result, and DeepSeek V3.2 carried the previous generation's.
- Four cells in the features comparison table had drifted from what the compared products now do.
- Four silent limitations are now written down: a Chat Completions message
nameis accepted and ignored, theverbose_jsontranscription decoder fields are placeholders rather than measurements, a bad batch input file is reported per line rather than refusing the batch, and the moderations input cap is the gateway's own rather than a vendor rule the docs implied it was.
v1.17.0 – 2026-09-08 – Your Own Models, Your Own Tenants, Your Own Spend¶
At a glance
- Your own model endpoints — Amazon SageMaker AI and Amazon Bedrock Marketplace endpoints join the model list and answer on the same routes, streaming included; one scaled to zero is woken rather than refused.
- A new dialect — clients written for the Ollama API reach Amazon Bedrock unchanged, chat, generate, embed and the discovery verbs.
- Per-tenant API keys — one deployment serves several tenants, each scoped to the models and endpoints it may use, revocable on its own and free to bring its own AWS account.
- Usage and spend from the API — the OpenAI Administration API reports what was used and what it cost, per model, per endpoint and per user. Off by default: it needs
USAGE_APIandCLOUDWATCH_METRICS. - WebRTC realtime —
POST /v1/realtime/callsterminates a browser's media session in the gateway, with no WebSocket relay in between. Off by default: it needsREALTIME_WEBRTC_ENABLED.
This release widens what a deployment can serve, who it can serve it to, and what it can tell you about the bill. Models you host yourself: Amazon SageMaker AI endpoints and Amazon Bedrock Marketplace model endpoints you have already deployed join the model list and answer on the same routes as everything else — a fine-tuned model, a model no serverless catalogue carries, or one you scale to zero between requests, reached with the same client code. A new dialect: clients written for the Ollama API, and the many tools that speak only that, now reach Amazon Bedrock unchanged. Per-tenant API keys: one deployment serves several tenants, each scoped to the models and endpoints it may use, revocable on its own, and free to bring its own AWS account. Usage and spend, from the API: the OpenAI Administration API reports what was used and what it cost — per model, per endpoint and, where a deployment identifies callers, per user — read from the metrics the gateway itself writes rather than estimated. Speaking to it, live: the Realtime API gained the WebRTC transport, so a browser negotiates its own media session and the audio path terminates in the gateway, with no WebSocket relay in between.
New IAM Permissions, All Optional
Nothing new is required. A deployment that enables none of the features below needs the policy it already has; the Terraform module grants each set when its feature is turned on. See IAM Permissions for the statements in full, and note that a denial now names the action it needs in the server log rather than leaving you to guess it.
- Your own SageMaker AI endpoints —
sagemaker:CallWithBearerTokenandsagemaker:InvokeEndpointon the endpoint ARNs you serve, for Amazon SageMaker AI endpoints. - Your own Marketplace model endpoints —
bedrock:ListMarketplaceModelEndpointsandbedrock:GetMarketplaceModelEndpointto discover Marketplace model endpoints,sagemaker:InvokeEndpointandsagemaker:InvokeEndpointWithResponseStream(called via Bedrock on the server's behalf) to invoke them, and — for a deployment with per-user roles —bedrock:InvokeModelonarn:aws:bedrock:*:ACCOUNT_ID:marketplace/model-endpoint/*, which afoundation-modelwildcard does not cover. - Shared records — the
dynamodbitem actions on one table, for per-tenant API keys or the shared model list. - Tenants —
ssm:PutParameterandssm:GetParameterto deliver each tenant its key once,kms:Encrypt,kms:Decryptandkms:GenerateDataKeyon the key named byTENANT_KEY_SSM_KMS_KEY_IDwhen you set one, andsts:AssumeRoleon exactly the roles tenants declare when they bring their own AWS account. - Usage reporting —
cloudwatch:GetMetricDataandcloudwatch:ListMetricsfor the Administration API. - Per-user cost attribution roles — one addition to an existing feature: the role the gateway assumes per end user now needs
s3:GetObject(andkms:Decryptthrough S3) on every bucket a request can reference by URI, because Amazon Bedrock reads a video or document at ans3Locationwith the invoking identity. Without it, a request that carries stored media answers403under that role while the same request under the task role succeeds. The module grants it; a hand-written role needs the statements in Per-User Cost Attribution.
Behavior Changes
Review these before upgrading — they may change what existing clients observe, and what AWS charges you:
- Three responses now name the model that actually served the request, rather than echoing the string the caller sent. The Anthropic Messages response (
/anthropic/v1/messages, streaming and not), the batch object returned byGET /v1/batches/{id}, and themodelfield inside Anthropic batch result lines — so a request namingclaude-opus-5now comes back naminganthropic.claude-opus-5. This matches what the OpenAI chat, Responses, embeddings and video responses already did, and affects existing callers who use aliases, not only callers using the new wildcard patterns. - A price change for the OpenAI GPT-5.6 family. GPT-5.6 Sol, Terra and Luna are now served by Amazon Bedrock Mantle by default, so their
web_searchandcode_interpretertools work with no configuration — Bedrock serves those on Mantle alone. Mantle has no cross-region inference profiles, so these models stop riding the Global profile and its discount and pay the In-Region rate: exactly 10% more per token, on input, output, cached and long-context rates alike. Per million input / output tokens that is $4.40 / $22.00 for Sol, $2.20 / $13.20 for Terra and $0.22 / $1.32 for Luna, against $4.00 / $20.00, $2.00 / $12.00 and $0.20 / $1.20 before. A deployment that had already disabledAWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBALpays what it paid. SettingAWS_BEDROCK_MANTLE_PREFERRED_MODELSto an empty value restores the previous routing, and the previous price, for every dual-homed model. - A guardrail is never silently skipped on a Mantle-served request. Amazon Bedrock Guardrails do not apply to Amazon Bedrock Mantle, so a model served there under a configured guardrail would answer a caller who asked to be guarded with an unguarded answer, and nothing would report it. Both halves of that are now refused. A deployment that routes a dual-homed model to Mantle under a guardrail — including through a
MODEL_ALIASESguardrail — stops at startup, naming the routed models and the way out; becauseAWS_BEDROCK_MANTLE_PREFERRED_MODELSnow has a default, a guardrailed deployment meets this without having changed anything, and clearing that setting restores both the guardrail and the previous routing. A model that only Mantle serves has no such way out: with a guardrail configured — deployment-wide, throughMODEL_ALIASES, or named on the request itself — every request to it answers400, whatever that setting holds. - The same rule now covers your own SageMaker AI endpoints. A request to a SageMaker-served model under a configured guardrail answers
400instead of being served unfiltered, and aMODEL_ALIASESguardrail naming one stops startup. A SageMaker endpoint with no capacity to wake answers503(retry later), not400. - Mantle routing and the per-user role's identity requirement are refused together. A Mantle request is signed with the server's own credentials, so
AWS_BEDROCK_USER_ROLE_REQUIRE_IDENTITYcan never apply to a model routed there: a deployment setting both withAWS_BEDROCK_MANTLE_PREFERRED_MODELSnon-empty stops at startup, and a Mantle request without an end-user identity answers400. - The Terraform module requires Terraform or OpenTofu 1.9. An existing deployment has to be on that version before it can upgrade the module.
- The Terraform module now scales the service on CPU, where before it never scaled at all.
autoscaling_cpu_target_percentdefaults to 70 instead of creating no scaling policy, so a fleet that has always held one task per availability zone now grows under load up toautoscaling_max_capacity— five times the minimum unless you set it. That raises what the Marketplace license can bill: a 3-AZ deployment metered at $0.10 per container-hour moves from a flat ~$216/month to a range whose ceiling is five times that. The gateway spends most of its time awaiting AWS, so CPU mainly answers the audio and video transcoding; setautoscaling_alb_target_requests_per_targetto scale on request volume instead, orautoscaling_cpu_target_percent = nullto hold the fleet at a fixed size. A WebRTC media deployment is unaffected: it pins the capacity to one task, and no policy is created when the minimum and maximum are equal. - A load balancer with no certificate now says so.
alb_enabledwithoutalb_domain_namein a public zone or analb_certificate_arnserves the API over plain HTTP, which is a supported way to stand up a trial. The plan and apply now carry a warning naming the two variables that turn on HTTPS, rather than leaving it to be noticed. It is a warning, not an error: nothing about the deployment changes. - The Terraform module turns its CloudWatch alarms on for a deployment that named a topic.
alarms_enableddefaults to whethersns_topic_arnis set, rather than tofalse: a deployment that gave a topic and never set the flag had no alarms and no notifications, which is the one configuration that cannot have been intended. It now creates the five it was configured for — the four ECS service alarms of the underlying module, plus one onERRORandCRITICALlog lines — each with its own CloudWatch charge. Settingalarms_enabled = falsekeeps the previous behaviour, and setting ittruewith no topic creates the alarms with nothing to notify. - A configuration that would silently drop the compliance VPC endpoints is refused at plan time.
compliance_vpc_endpoints_enabledandguardduty_vpc_endpoint_enabledneed a private application subnet to place their network interface in. Settingnat_gateways_allowed = falsewhile internet access is still required — which it is by default — makes those subnets public, and the endpoints were then simply not created, with nothing said. That combination now fails the plan. A deployment already running it was not getting the endpoints it had asked for: setnat_gateways_allowed = true, or leave it unset, or turn the endpoints off. - Input token counting is unavailable for the GPT-5.6 family.
POST /v1/responses/input_tokensanswers400for the models Mantle serves, which now includes that family. ClearingAWS_BEDROCK_MANTLE_PREFERRED_MODELSbrings it back. - GPT-5.6 usage is reported and billed under Bedrock Mantle, attributed by project rather than by IAM principal, and runs on Mantle's own throughput quotas. Batch inference, prompt caching and response IDs stored before the upgrade are unaffected.
- The community container image moves from Alpine to a standard Debian base. glibc is what the WebRTC media stack needs —
aiortcpublishes no musl wheel for any version — so the community image now ships every capability the commercial one does. It declarespythonas the image entry point and runs as an unprivileged user with uid/gid 65532 instead of 1000. The API, its endpoints and the audio formats it accepts are unchanged, and deployments through the Terraform module are unaffected — the task definition already pinned 65532. Two things change for anyone running the image by hand: arguments after the image name are passed to the Python interpreter rather than run as a program, and a mounted~/.awsneeds--user "$(id -u):$(id -g)" -e HOME=/home/nonroot— or--userns=keep-id:uid=65532,gid=65532on rootless Podman — to stay readable. It is also a larger image, built from the distribution's own packages; the hardened, minimal one is the AWS Marketplace edition. See Local Development.
Your Own Model Endpoints¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| stdapi.ai | Amazon SageMaker AI endpoints – the endpoints you already run join the model list under the names you pick and answer on the same routes as every other model, streaming included: a fine-tuned model, or one no serverless catalogue carries, reached with unchanged client code | |
| stdapi.ai | Endpoints scaled to zero – an endpoint with no running capacity is woken and waited for rather than refused; one with no capacity to wake answers 503 (retry later), not 400 | |
| stdapi.ai | Amazon Bedrock Marketplace model endpoints – the endpoints you have deployed are discovered and served like any other model, invoked through Amazon Bedrock on the server's behalf, and one being updated stays served |
New APIs¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
POST /ollama/api/chat – conversational responses with tools, images and thinking, streamed as the newline-delimited JSON Ollama clients expect, carrying the timings they read | ||
POST /ollama/api/generate – a response for a single prompt, streamed or whole; an empty prompt, here or on /api/chat, is the load/unload no-op Ollama defines rather than a 400 | ||
POST /ollama/api/embed and POST /ollama/api/embeddings – vector embeddings for one or several inputs, including the legacy single-prompt route older clients still call | ||
/ollama/api/tags, /ollama/api/show, /ollama/api/ps, /ollama/api/version and /ollama/api/pull – the discovery verbs an Ollama tool calls before anything else. The verbs that write to a local model store — create, copy, push, delete — are refused, since this deployment stores no models of its own. The /api/blobs/{digest} upload joined that refused set in v1.18.0, which also answers Ollama is running on the base URL for a client's connection test | ||
GET /v1/organization/usage/* – tokens, images, characters and seconds per bucket, per model, per endpoint and — where a deployment identifies callers — per user, read from the metrics the gateway itself writes. An endpoint that measures nothing answers empty buckets rather than a 404. Off by default: it needs USAGE_API and CLOUDWATCH_METRICS, and grouping by user needs CLOUDWATCH_METRICS_USER_DIMENSION as well | ||
GET /v1/organization/costs – what AWS bills this deployment for serving those requests, over the same buckets and the same filters. Needs COST_TRACKING on top of the two settings above | ||
POST /v1/realtime/calls – the WebRTC transport: a browser posts its own SDP offer, negotiates a media session directly with the gateway and the audio path terminates there, with no WebSocket relay in between. An offer whose ICE candidates are all private, link-local or mDNS names is refused with invalid_offer; a deployment on a private network sets REALTIME_WEBRTC_ALLOW_PRIVATE_CANDIDATES. Off by default: it needs REALTIME_WEBRTC_ENABLED, and the task must be directly reachable over UDP — behind an HTTP-only load balancer the SDP exchange succeeds but the call carries no audio |
Multi-Tenancy¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| stdapi.ai | Per-tenant API keys – issue a key per tenant, scope it to the models and endpoints that tenant may use, and revoke it without touching anyone else. Files, vector stores, batches, conversations and stored responses stay deployment-wide with no tenant ownership — a documented limitation of this release | |
| stdapi.ai | Key delivery that skips Terraform state – the server mints each tenant's secret and delivers it once through a parameter, encrypted with the key named by TENANT_KEY_SSM_KMS_KEY_ID — the deployment's own KMS key when the Terraform module provisions it — so reading it takes a grant on that key and not merely permission on the parameter path | |
| stdapi.ai | Tenant AWS credentials – a tenant may bring its own AWS account: its requests then run under its own role, its own Amazon Bedrock quota and its own bill |
Model Discovery¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| stdapi.ai | Wildcard model names – any request that names a model may name a pattern instead — claude-sonnet-*, say — and the most recently released match serves it; an ambiguous tie is refused rather than guessed | |
| stdapi.ai | GET /search_models – a model= filter lists everything a pattern matches, newest first, so a pattern can be checked before anything relies on it | |
| stdapi.ai | Shared model list – a fleet discovers models once and shares the result, instead of once per task, and the refresh stays off the request path |
API Improvements¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| Diarized transcripts streamed – a diarized transcription is returned as speaker-labelled segments as they are recognized, instead of being refused | ||
| Fused multimodal embeddings – text and an image embedded together into one vector, on both embed routes | ||
| stdapi.ai | Cohere Embed v3 listed as image-capable – Bedrock publishes cohere.embed-english-v3 and cohere.embed-multilingual-v3 as text-only although both embed an image and are billed for it; the catalogue now declares the modalities the listing omits, so /v1/models, /search_models and the Models page agree with what is served |
Platform Features¶
| Feature | Description |
|---|---|
| Public Models page | Browse, filter and compare every model the gateway serves, with each fact resolved from the source that states it |
| Denials name the permission they need | An AWS authorization failure is logged with the action that was denied, instead of leaving the operator to infer it from the call that failed |
| MCP server identity | The published container images declare the MCP server they serve, so a registry and an MCP client can identify it |
| WebRTC media path provisioned | Off by default: it needs REALTIME_WEBRTC_ENABLED on the server and realtime_webrtc_media_enabled on the Terraform module. In media mode the module provisions the whole UDP path on the VPC it creates — the matching egress, and the application-subnet network ACL entries the media needs; bring your own subnet_ids and those entries stay yours to widen. The module's realtime_webrtc_stun_server defaults to a Google-operated public STUN server, which the task reaches over an egress rule open to 0.0.0.0/0 on that port — the one place this deployment talks to a third party. Point it at a STUN server you run to keep the flow inside your account |
| Every setting reachable from Terraform | The module carries every setting the server accepts, so an existing table, tenant key prefix or KMS key can be named instead of created, and the grants follow what is named |
| Per-user attribution role, built by the module | aws_bedrock_user_role_create creates the role the gateway assumes per end user — its trust policy and the invoke permissions AWS authorizes against the caller included — instead of leaving it to be written by hand and named by ARN; aws_bedrock_user_role_arn still takes an existing one and wins where both are set. Off by default, as the feature itself is |
| Scales on CPU out of the box | A deployment that sets no autoscaling variable now scales on CPU at 70%, between one task per Availability Zone and five times that — size the ceiling before it surprises the bill |
| Home Assistant voice, end to end | The deployment sample now configures the conversation agent as well as speech, so a spoken command is understood and acted on rather than only transcribed and spoken back |
Fixes¶
- Images reach the models that read them: the Qwen vision-language line was published by the open-weight family, which advertises text only, so an attached image was refused by models that in fact read it. The line now declares image input and is excluded from the open-weight matcher, so iteration order can no longer decide what a model is published as.
- The Anthropic MCP connector no longer blocks a request: an
mcp_serversrequest was rejected outright and themcp_tool_useandmcp_tool_resultblocks it produces had no accepted shape, so a client using the connector could not call the API at all. The field and the blocks are accepted, and the request is served without the connector rather than refused — it names an extra tool source, it is not the request itself. - An embedded image is billed for what AWS charges: Cohere Embed v3 meters an image as its own billed unit and reports no token count for it, so image embeddings were attributed at zero cost while AWS billed for them. The images the response reports are now recorded against the model's image rate; later Embed versions bill images inside the tokens and are unaffected.
- A model both catalogues name stays reachable through bedrock-runtime: naming a dual-homed model in
AWS_BEDROCK_MANTLE_PREFERRED_MODELSpublished the preferred entry under that identifier and dropped the other, so the model was reported impossible to batch — batch inference runs on bedrock-runtime alone. The displaced bedrock-runtime entry is kept off-catalogue and read by everything that must reach that endpoint, while the published catalogue keeps naming the preferred entry. It now matters to every deployment, that setting having gained a default. - Listings report the time they are ordered by: batch and completion listings paged in identifier order while reporting a creation time recorded from another clock, so the sequence a client saw did not match the timestamps it was given and a page boundary could revisit or skip an entry. The reported time is now read from the identifier itself, and a batch listing finds the newest batches whatever the bucket holds, including after a dense burst.
- A one-token answer is served rather than refused: the Anthropic Messages API and Chat Completions both accept a budget of a single token — a client probing a model sends exactly that — while the Responses transport a Bedrock Mantle model is served over refuses anything under 16, so such a request came back as a
400namingmax_output_tokens, a field the caller never sent. The budget is now raised to what the transport accepts. - An unknown model is a
404when counting input tokens:POST /v1/responses/input_tokensanswered400for a model that does not exist, where the official API answers404model_not_found— a client branching on the status read a malformed request instead of a missing model. The image routes keep their400, which is what upstream answers there. Token counting remains unavailable for models served by Bedrock Mantle or a model endpoint, which is a different400. - A deployment on your own IPv4-only subnets stops being told it is dual-stack: the VPC module read "has IPv6" from a subnet attribute AWS reports as an empty string rather than as absent, so every
subnet_idsdeployment looked dual-stack whatever its subnets carried. The load balancer was then asked fordualstack, anAAAArecord was published for an address nothing answered on, and the server was told to bind::. The module now reads the subnet's actual IPv6 block. - A model's release date is when it launched, not when your region got it: a region that opens a model months after launch reports that rollout as the model's own start-of-life, and whichever region happened to be listed first decided the published date — Claude Haiku 4.5 reads as 2026-01-06 in the seven regions it expanded into and 2025-10-15 in the twenty-four that had it at launch. The earliest date any region publishes is now the one reported, by
createdin the OpenAI model list, by the Anthropic and Ollama model routes, and on the Models page. - The Models page states only what a live source backs: a benchmark row is matched to the model release it was measured on, a source that publishes no row now fails the build instead of leaving stale figures in place, facts no longer leak between variants of a family, and a lifecycle date follows one stated rule — the Bedrock API is the reference, and a model card fills a date only where the API states none.
- A deployment on your own subnets can be planned again: the VPC module this one builds on fed a set to a function that takes only lists, so every plan that creates a VPC failed outright — including for a deployment that had changed nothing, since the constraint resolves to the newest release. Fixed in
JGoutin/terraform-aws-vpcv1.6.1. - A deployment on your own subnets can be planned again, twice over: the module's IPv6 out-of-band egress rule read whether the subnets carry IPv6 before checking that the WebRTC media path was even on. On
subnet_idsthat answer comes from the subnets themselves and is unknown while planning, so the plan failed outright — with the rule set empty and WebRTC off, which is every such deployment. - A module that names your own VPC no longer deadlocks the plan: the length limit
name_prefixcarries when a load balancer is enabled was asserted on the variable, which tied it toalb_enabledand through that tosubnet_ids. Passing the module'sname_prefixoutput to a VPC and that VPC's subnets back — what the output exists for — closed a dependency cycle. The limit is now checked on the load balancer itself and still refuses too long a prefix while planning. - The documentation pages load a current Swagger UI:
swagger-ui-dist5.32.15, which scopes its HTML sanitiser to itself instead of mutating the page's global one. Both pins moved on in v1.18.0, toswagger-ui-dist5.33.0, which lays out only the part of a large specification that is on screen, andredoc2.5.4, which carries accessibility fixes and patched bundled dependencies.
v1.16.0 – 2026-08-21 – Conversations, Batches, Vector Stores, Realtime Speech & Per-User Identity (with v1.16.1 maintenance update, 2026-08-25)¶
At a glance
- Conversations and batches — Conversations keep a thread server-side so a client continues it by id; the OpenAI Batch API and the Anthropic Message Batches API run large request sets at the discounted batch price.
- Vector stores — index and search files by meaning, or address a knowledge base you already run, with a model reaching either for itself through
file_search. - Realtime speech — the Realtime API holds a spoken conversation over one WebSocket, with a transcript of both sides, turn detection and barge-in.
- Speech and audio — 100,000-character synthesis spoken as it is produced, live transcription needing no bucket, and Amazon Nova Sonic as a speech-to-text backend.
- Identity per caller — Amazon Cognito tokens alongside or instead of the API key, published discovery, and per-user cost attribution read from the AWS invoice. New IAM permissions are required, and five features stay inert until you create the resource they need.
This release adds four API surfaces and finishes the speech story. New APIs: Conversations keep a thread server-side, so a client continues it by id instead of resending the history; the OpenAI Batch API and Anthropic Message Batches API run large request sets asynchronously at the discounted batch price; Vector Stores index and search files by meaning — or address a knowledge base you already run — with a model reaching either kind for itself through file_search; and the Realtime API holds a spoken conversation over one WebSocket. Speech: 100,000-character synthesis spoken as it is produced, live transcription needing no bucket, and Amazon Nova Sonic as the lowest-cost speech-to-text backend here. Identity per caller: Amazon Cognito tokens alongside or instead of the API key, published discovery so an agent authenticates itself, and per-user cost attribution reporting each end user's spend from the AWS invoice rather than an estimate.
New Required IAM Permissions
v1.16.0 adds one action every deployment needs, a handful that belong to statements you may already grant, and one statement per optional feature. See IAM Permissions for the policies in full.
Enough of them together that they no longer fit one policy: IAM caps a customer managed policy at 6,144 characters, and a deployment enabling most of these exceeds it. Attach several policies to the role rather than widening actions to save room — the Terraform module now ships two, one for Amazon Bedrock and one for the supporting services, and does that for you.
Required on upgrade, whatever the deployment does:
bedrock:InvokeModelWithBidirectionalStream— serves every model invoked over a two-way connection — the Realtime API, and Amazon Nova Sonic transcription and translation. It belongs to the core Bedrock policy; no route-specific action exists for any of them.
Add to a statement you already grant, if the deployment uses that feature:
bedrock:UpdateSession, on the session storage statement — the conversation metadata update (POST /v1/conversations/{id}) and nothing else. The rest of the Conversations API uses the actions stored responses already require.bedrock-mantle:CountTokens, on the Bedrock Mantle statement — counts the tokens of a Mantle-served model on/anthropic/v1/messages/count_tokens, since Amazon Bedrock's ownCountTokenstakes Anthropic models only. Needed by any deployment serving Mantle models, which is the default.transcribe:StartStreamTranscription, on the speech-to-text statement — servesstream=trueon/v1/audio/transcriptions. Add it with the upgrade: without it a streamed request that names its language answers503feature_unavailable, and the server log names the permission. It stages nothing, so a deployment with no bucket at all grants this one alone.polly:StartSpeechSynthesisStream,polly:StartSpeechSynthesisTaskandpolly:GetSpeechSynthesisTask, pluss3:PutObject,s3:GetObjectands3:DeleteObjecton each bucket serving an Amazon Polly Region, on the text-to-speech statement — they serve input above 3,000 characters and nothing else. With a bucket configured and these missing, long requests are accepted and then fail on the permission: grant the whole set, or leave the bucket unconfigured and keep the 3,000-character answer.translate:ListLanguages, on the translation statement — read once at startup so an unsupported language pair is refused before the audio is transcribed. Genuinely optional: without it the check stays off and translation still works, reporting the unsupported pair once the translation call itself fails.
New statements, one per optional feature:
- Vector stores —
s3vectors:CreateIndex,DeleteIndex,PutVectors,GetVectors,QueryVectorsandDeleteVectors, scoped to your vector bucket and its indexes. No bucket-level create or delete is granted: the gateway creates and deletes the indexes inside the bucket, never the bucket. - Knowledge base vector stores —
bedrock:GetKnowledgeBase,Retrieve,ListDataSources,IngestKnowledgeBaseDocuments,ListKnowledgeBaseDocuments,GetKnowledgeBaseDocumentsandDeleteKnowledgeBaseDocuments, one statement per allowlisted knowledge base ARN.bedrock:ListKnowledgeBasesis deliberately not granted and not needed — the server only ever addresses the identifiers it was given. - Batch inference —
bedrock:CreateModelInvocationJob,GetModelInvocationJobandStopModelInvocationJobon the server's role, plusiam:PassRoleconditioned onbedrock.amazonaws.com. The service role Amazon Bedrock assumes carries its own policy:s3:GetObject,s3:PutObjectands3:ListBucketon the batch prefix, andbedrock:InvokeModelon the models you batch. - Per-user cost attribution —
sts:AssumeRoleandsts:TagSessionon the server's role and in the end user role's trust policy (both actions: withoutTagSession, every tagged session is denied), andbedrock:InvokeModel,bedrock:InvokeModelWithResponseStreamandbedrock:ApplyGuardrailon the end user role itself, since AWS authorizes those against the caller of the invocation. - Web search —
bedrock-websearch:InvokeSearchandbedrock-websearch:InvokeFetch, plusbedrock-websearch:ExternalWebAccessonly where a request may reach the open internet. Leaving that last one out is what keeps every search inside the AWS boundary. A missing web-search permission produces no error and no server log entry: the model answers without having searched, so check these before suspecting the model. - Transcription output encryption —
kms:GenerateDataKeyandkms:Decrypton the key named byAWS_TRANSCRIBE_OUTPUT_ENCRYPTION_KEY_ARN, in the key policy as well as on the role. - Durable vector store indexing —
sqs:SendMessage,sqs:ReceiveMessage,sqs:DeleteMessage,sqs:ChangeMessageVisibilityandsqs:GetQueueAttributes, on the single queue named byAWS_SQS_VECTOR_STORE_QUEUE_URLand never on*. Needed only if you configure that queue; leave it unset and nothing here applies. No queue is ever created, deleted or reconfigured, so none of those actions is granted.
Five features stay inert until you create the resource they need
Nothing in this release is breaking — but these five answer 503, or stay off, until the resource exists in your own account:
- Batches — an IAM service role Amazon Bedrock assumes to read the requests and write the results (
AWS_BEDROCK_BATCH_ROLE_ARN), plus the bucket it reads and writes. - Vector stores — an Amazon S3 vector bucket you create yourself, and the Region it lives in.
- Knowledge base vector stores — an allowlist of the knowledge bases this deployment may address (
AWS_BEDROCK_KNOWLEDGE_BASE_IDS), empty by default. One that is not on it answers exactly as a store that does not exist, so the setting cannot be probed for what a deployment holds. - Cognito authentication — a user pool and its app clients (
AWS_COGNITO_USER_POOL_ID); until then the API key remains the only method, exactly as before. - Per-user cost attribution — a role for the end user sessions (
AWS_BEDROCK_USER_ROLE_ARN); off by default, and every call keeps being billed to the deployment's own identity until it is set.
Conversations, the Realtime API and streamed transcription need no new resource. Long speech input needs a bucket for the serving region, which is the same one the rest of the gateway already uses — except on generative voices, which speak up to 20,000 characters without one.
Behavior Changes
Review these before upgrading — they may change what existing clients or dashboards observe:
- A missing deployment permission is no longer reported as the caller's. An
AccessDeniedExceptionon the gateway's own AWS calls reached clients as403 permission_error— which every OpenAI and Anthropic SDK reads as their key being refused. Every route now answers503feature_unavailable, with the server log naming the operation, model and permission. Clients matching403for a backend permission error should match503/feature_unavailableinstead; a403now means only that per-user attribution is on and that end user's role was denied. - Built-in web search now appears in usage and cost reporting. Queries were recorded as nothing at all, so a measured turn under-reported its cost by 58%. Nothing AWS charges changed; what the gateway reports does. Web access is also an operator setting now (
AWS_BEDROCK_EXTERNAL_WEB_ACCESS), defaulting to the previous behaviour. - A request that would be answered without what it asked for is refused. A
/v1/responsesweb_searchrestricting its sources (filters.allowed_domains,user_location) was accepted and dropped, so answers came back sourced from domains the caller had excluded. Now a400on models that cannot serve it; Bedrock Mantle models receive the options unchanged. The same rule governs file search filters and score thresholds. - Two output-shaping hints that returned
400now succeed.predictionandverbosityon chat completions are accepted and dropped, as the Responses surface already did;truncation="disabled"is likewise accepted, whiletruncation="auto"is still refused. /v1/responsesforwards undeclared request fields to the model, as chat completions and messages already did, so the backend may refuse one it does not recognise. Conversely, client-side control fields no provider treats as parameters (LiteLLM'sdrop_paramsamong them) are dropped rather than forwarded. Both are governed byEXTRA_MODEL_PARAMS_DENYLISTandEXTRA_MODEL_PARAMS_DROP_ALL.- Attachments are measured against what the model actually accepts. The old guard compared raw bytes where the backend enforces base64 length, so it was ~33% too permissive. Oversized attachments are now staged and referenced where the model reads from storage, or refused with
413naming the size it accepts. Smaller attachments are unaffected — see Attachment Size. - The server's own connections follow the proxy environment.
HTTPS_PROXY,HTTP_PROXYandNO_PROXYwere honoured by the AWS SDK and ignored by everything else, so a proxied deployment saw no Bedrock Mantle models. Two connections deliberately still bypass it: container metadata, and the fetch of a caller-supplied URL, where a proxy would defeat address validation. See proxied deployments. - A declared upload checksum is now verified. The value was stored and never looked at, so a corrupted upload completed like a clean one. It covers the file's contents, not the storage layer's multipart identifier — declaring the latter is now refused.
- An unknown model name answers with a sentence, not the catalogue. The
404body carried every served identifier, roughly 2,500 characters. Clients that parsed it for a model list should call/v1/models. - Bedrock Mantle is only probed in the Regions that serve it, so a deployment listing others no longer warns at every start. An explicit
AWS_BEDROCK_MANTLE_REGIONSlist is still used exactly as given. - The container health probe's command changed. Deployments that re-declare the probe instead of running the image's own — an ECS task definition among them — should take the command from the image.
New APIs¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/conversations – create, retrieve, update and delete a conversation, list and manage its items, and continue it from the Responses API with the conversation parameter | ||
/v1/vector_stores – attach files, follow the indexing as it progresses, then search by meaning with attribute filters and per-passage scores | ||
/v1/vector_stores – address an Amazon Bedrock knowledge base you already run as a vector store, Bedrock managed or customer-managed: search it, attach, list, read and delete documents. Allowlisted per knowledge base, never created or deleted here | ||
file_search on /v1/responses – a chat model answers from the stores you name, reporting the searches it ran and citing a file_citation per file it drew on | ||
/v1/batches – run a JSONL file of chat completion or embedding requests asynchronously at the batch price: submit, poll, cancel, read the result files | ||
/anthropic/v1/messages/batches – the same asynchronous, batch-priced run for the Messages API, results streamed back as JSONL; each request may name its own model, up to eight per batch | ||
WS /v1/realtime – a live speech-to-speech session over one WebSocket, with a transcript of both sides, server-side turn detection or manual turns, barge-in, and G.711 for telephony | ||
POST /v1/realtime/client_secrets – mint a short-lived, browser-safe credential carrying a session configuration; signed and stateless, so any instance verifies one minted by any other |
Limits worth knowing before building on these
Realtime: a session lasts at most 8 minutes and calls no tools, and a spoken answer is guardrail-checked once complete, so a blocked one may already have been partly heard (coverage). WebRTC and SIP are not served in this release — put LiveKit Agents or Pipecat in front for a browser media path or a phone line. Superseded in v1.17.0, which added the gateway-terminated WebRTC transport as an operator opt-in; SIP is still never terminated here. Tool calling was added in v1.18.0, so a session does call tools. The compatibility table lists every event the session does not emit.
Knowledge-base stores address a knowledge base that already exists and refuse, naming why, anything that would reshape it — creating, deleting, renaming, expiry, chunking strategy, attribute rewrites and the file-batch routes. Attaching needs a custom data source. Retrieval scores are reported as the backend states them rather than rescaled into similarities, and unknown values are reported unknown rather than invented. See Knowledge Base Stores.
Speech & Audio¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/audio/speech – up to 100,000 billed characters per request, 24× the upstream 4,096, with no API change and no new request field | ||
/v1/audio/speech – long input is spoken as it is synthesized instead of after a whole job finishes; generative voices reach 20,000 characters with no bucket at all, and each request takes whichever path can serve it, so long input is no longer tied to one voice | ||
/v1/audio/transcriptions – naming Amazon Nova Sonic transcribes at the lowest cost available here, streamed as it is recognized; json and text only, up to 10 minutes, no timestamps. No existing request is re-routed | ||
/v1/audio/translations – Amazon Nova Sonic translates speech to English itself, in one request | ||
/v1/audio/transcriptions – stream=true returns each phrase as it is recognized, whenever the request names the language to expect; needs no bucket. gpt-live-transcribe is now an alias, and requests naming no language are unchanged unless AWS_TRANSCRIBE_STREAM_LANGUAGES says which to expect | ||
/v1/audio/transcriptions – per-language custom vocabularies and language models, so a request identifying between several languages can apply the right resources to each one instead of being refused. Accepted only where the backend would use them: alongside a single fixed language, where they would apply to nothing, they are still refused | ||
| stdapi.ai | AWS_TRANSCRIBE_OUTPUT_ENCRYPTION_KEY_ARN – encrypt a transcription's output with a key you name rather than the bucket's own. The job's request identifiers travel as the encryption context, so a key policy can be scoped to this workload instead of to the whole bucket | |
/v1/audio/translations – the supported language pairs are read once at startup and checked before the call, so a pair that cannot be served is named as the request problem it is instead of surfacing as a failure after the audio was transcribed. The permission that reads them is optional: without it the check stays off and everything else works |
Identity & Cost Attribution¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| stdapi.ai | Amazon Cognito user pool tokens – accept access tokens instead of, or alongside, the API key, so each caller reaches the API with their own credential; validated in-process against the pool's published keys, with no AWS call on the request path | |
| stdapi.ai | AUTHENTICATION_MODE – assert the posture rather than infer it: the server refuses to start when the selected method is not configured, or when a configured method would be silently ignored | |
| stdapi.ai | Authentication discovery for agents – an OAuth 2.0 protected resource metadata document, pointed at by every unauthorized response, so an MCP client finds the authorization server and the scope it needs without being configured for this deployment; published only once an authorization server is declared | |
| stdapi.ai | Per-user cost attribution – model calls issued under a short-lived role session tagged with the caller, so AWS reports each end user's spend in Cost Explorer and the Cost and Usage Report, from the invoice rather than an estimate. Off by default; a deployment can also require every call to name its end user rather than bill it to the deployment | |
| stdapi.ai | Vector store cost reporting – a search against a Bedrock-managed knowledge base is recorded and priced like every other billed unit; what cannot be accounted for is stated rather than approximated |
Platform Features¶
| Feature | Description |
|---|---|
| Aliases that carry configuration | A MODEL_ALIASES entry may map a public name to the target model plus the service tier, guardrail, metadata and extra parameters applied to requests naming it, so one model is published under several names with different policies. The plain-string form is unchanged, and a malformed alias stops startup naming itself rather than failing once per request |
| Attachment size policy | On the multimodal routes served by Amazon Bedrock, an attachment is measured before the request is built and travels inline or by reference according to the limits each model class declares. Staging is per model and per media kind: of the families measured, only the Amazon Nova families and TwelveLabs Pegasus accept a reference |
| Ephemeral secret signing key | REALTIME_CLIENT_SECRET_KEY signs the Realtime API's client secrets. A deployment with an API key already shares one and needs nothing; one running with no API key at all should set it, or a secret minted by one instance fails to verify on another |
| Dual-stack container listener | The image's bind address moved out of its command into GRANIAN_HOST, so a deployment that needs a dual-stack socket — an ECS service whose discovery record includes an AAAA record, for instance — sets one variable instead of replacing the whole command. The IPv4-only default is unchanged |
| Faster container health probe | The probe ships as a module of the application itself, byte-compiled with the rest of the package and covered by the linters and the test suite; it speaks HTTP over a socket rather than pulling in 123 modules per probe, cutting roughly 250 ms of import work per run in the community image and halving its peak memory |
| OpenAI Daybreak models | Daybreak Red (GPT-5.6 Cyber) and Daybreak Blue (GPT-5.6 Sol) are served and priced with the rest of the GPT-5.6 family, image input included. Both answer on the Responses API through Bedrock's next-generation inference endpoint, in US East (Ohio) only, and both are gated on enrollment with OpenAI's Daybreak programme — an account without it does not see them in the catalogue at all |
| Capability discovery | The model catalogue advertises what this release added — speech to speech, its transcription and translation, the search surfaces, and whether a model can be used with the Batch API — filterable over HTTP and through the same tool an agent reads before it calls anything. Web search is credited to every model that provides it, not only to the family the last release added it for |
| Durable vector store indexing | AWS_SQS_VECTOR_STORE_QUEUE_URL hands indexing to an Amazon SQS queue you create, so a file keeps being indexed — and finishes — when the server that accepted it is replaced. Off by default; needs a standard queue with a dead-letter queue and the durable indexing permissions (resilience) |
| End-to-end client coverage | The suite driving complete, unmodified third-party clients against a live gateway gains five: LiteLLM, Docling Serve's vision pipeline, the OpenAI Agents SDK (realtime voice, conversations, web search, vector-store retrieval), and LiveKit Agents and Pipecat running the exact WebRTC and telephony configurations this documentation prints. Deployment guides follow for LobeHub and RAGFlow |
Fixes¶
- Vector store durability: an attached file whose indexing was interrupted is now reported
failedwith alast_errorinstead of sittingin_progressfor ever; a detached file leaves every listing and search immediately and is gone only once its passages are, so a server replaced mid-delete no longer leaves a deleted document answering searches; indexing is bounded for the whole server rather than per request, so memory and embedding quota no longer scale with the number of callers; and deployments that would rather the work finished than reported can hand indexing to a queue - Work a request left running is finished before the server stops: a deployment, scale-in or Spot interruption used to drop temporary file cleanups, vector store indexing and live audio session releases with nothing in the logs. Shutdown now waits under
SHUTDOWN_DRAIN_TIMEOUT(10 seconds by default), settles whatever the deadline leaves, and counts it in thestopevent. Raise it together with your container runtime's kill delay, never one alone - Streamed responses run the work they scheduled: the drain was attached before the body produced a byte, so three leaks followed — a vector store searched only through streamed answers could expire mid-query, an expired store left a paid index behind, and a streamed transcription falling back to a job left its audio, transcript and job record behind on every request
- Realtime speaks the released vocabulary, not the beta one: item events were unparsable, the caller's transcript event was dropped for a missing field, and the item lifecycle clients wait on was never emitted. Barge-in did not work at all — the session refused truncation precisely while an answer was playing, the only moment it is ever sent. Truncate, retrieve and delete now work against the tracked conversation, a written turn is answered instead of timing out, and every answer reports the six response fields upstream always sends
- Batch API: listing a batch neither settled nor published it, so a client that only ever listed never had its usage recorded; every validation failure at submission was reported as an unsupported model, quota and role failures included; and the batch record was written only after the jobs started, leaving billable work running with nothing on disk to stop it
- Knowledge-base stores answer in this API's own words: attaching to a managed knowledge base always failed, listing its files always failed, deleted documents never left a listing, and internal bookkeeping reached the caller as attributes. A file refused because the corpus is maintained elsewhere is now told so in the store's terms, and a refused file is explained by the store that refused it rather than by one fixed sentence describing limits the caller never met
- Cost reporting matches what AWS charges: GPT-5.6 Luna was reported at five times its real cost and Terra a quarter over, since AWS repriced them on the model's own page rather than in the live price catalogue; a Global cross-Region call — the only way Luna, Sol and Terra are ever served — was costed at the In-Region rate, about 10% over; and moderation usage naming an alias matched no price at all. Every source page is now re-checked weekly
- Reasoning and web search reach the models that serve them: Amazon Nova 2 and DeepSeek V3 refused the token budget the Anthropic dialect requires, leaving extended thinking unreachable on that route while the identical ask worked elsewhere; and a
web_searchsent to a GPT-5.6 model resolved to its non-Mantle twin travelled as an ordinary function tool, so no search ran and nothing said so — now a400naming both ways to route the model to the endpoint that serves it - The Messages surface reports what an answer cost and why it stopped: refusals carry the policy category, the reasoning-token breakdown is reported, and service tier and per-TTL cache-creation counts are populated — most visibly on a batch, which claimed no tier while being billed as one
- Audio, embeddings and attachments are bounded correctly: inline audio was measured against raw bytes where the backend enforces the encoded length, so files between ~18.75 MB and 25 MB passed and were refused downstream; long text-to-speech now answers with the length a bucket-less deployment can honour; and models that embed one input per call no longer open a connection per chunk
- The interactive documentation pages render with no outbound access:
/docsand/redocpulled the icon, Swagger UI, ReDoc and a web font from three third parties — blank pages in an air-gapped VPC, and elsewhere a report to those hosts of who was reading this API and when, running whatever a floating major version tag resolved to that day. Both are now served whole from the image, pinned to exact releases verified by SHA-256 at build time, with upstream licences beside them - MCP tools return what they produce: every route answering with bytes was published as a tool an agent could call and then could not use. An image now arrives as an image and audio as audio, anything the protocol cannot carry arrives as a reference rather than failing, and the 4 MiB body cap that blocked image edits now follows
MAX_INPUT_FILE_SIZE - Addresses and listings are the ones this deployment serves: a custom route prefix still quoted default paths to
search_modelsand to video job polling, neither recoverable client-side; and a file'screated_atand its place in a listing came from two different clocks, so a multipart upload sat among older files reporting a later time. The Anthropic listing also answered oldest first, hiding every recent file, and now runs newest first as upstream does - Diagnostics name their cause: an unreachable Region rendered six different conditions as one identical sentence, and the slow startup beside it was the container metadata lookup retrying, unreported; a request abandoned mid-flight left its OpenTelemetry trace current, so later work was recorded under a closed request's trace id; and behind a proxy,
client_iprecorded the load balancer whateverPROXY_TRUSTED_HOSTSallowed - The API describes itself, not the service behind it: a synthesis limit credited to the service enforcing it, prices credited to their catalogue and a moderation route naming the engine underneath all shipped in the OpenAPI document and in the tool descriptions agents read before calling
Fixes & Maintenance (v1.16.1)¶
- Update
cryptographyto 50.0.1, rebuilt against OpenSSL 4.0.2
v1.15.0 – 2026-08-03 – Reliability, Performance & Feature Completeness¶
At a glance
- The largest correctness pass to date — three successive deep audits plus an independent full-branch review closed hundreds of fidelity gaps across all three API dialects, every fix pinned by tests.
- Performance — hot paths run in compiled native code and independent work runs in parallel, cutting the gateway's processing overhead.
- Explicit prompt caching — placement and lifetime control on the OpenAI dialect, plus fixes to the caching plumbing that already existed.
- Reasoning and prompt controls — operator reasoning controls, the Responses API
promptparameter from Amazon Bedrock Prompt Management, and native mid-conversation system messages on Claude 4.8+. - Verified with real clients — compatibility claims are backed by a test tier running complete, unmodified third-party client software against a live gateway. Behavior changes: unsupported parameters are now accepted and ignored, and a configured guardrail applies to every route.
This release focuses on making the whole gateway better rather than just bigger. Reliability and quality: the largest correctness pass to date — three successive deep audits plus an independent full-branch review closed hundreds of fidelity gaps across all three API dialects, every fix pinned by tests and the whole surface validated by real, unmodified client applications. Performance: hot paths now run in compiled native code and independent work in parallel, measurably cutting the gateway's processing overhead. Feature completeness: existing capabilities are rounded out end to end — explicit prompt caching, operator reasoning controls, the Responses API prompt parameter from Amazon Bedrock Prompt Management, native mid-conversation system messages on Claude 4.8+, richer speech and transcription (Polly speech marks, Transcribe/Translate extras, generic Converse speech-to-text), Cohere embedding_types and Rerank v1 structured documents, guardrail enforcement on every route, and an inline guardrail-checks moderation backend.
Behavior Changes
Review these before upgrading — they may change what existing clients observe:
- Unsupported parameters are accepted and ignored, not rejected. Parameters the AWS backends cannot honour (e.g.
known_speaker_*,partial_images, unsupported imagequality/style, programmatic tool calling) now behave like they do on OpenAI: the request succeeds, the parameter is dropped, and a warning is recorded in the request log. Requests that returned400on v1.14 may now succeed. - Impossible combinations are now clean
400s instead of silent degradation: subtitle or diarized formats withstream=true, contradictory Amazon Transcribe settings, and web-search filters Nova grounding cannot apply are rejected with actionable messages. - Speech output quality:
wav/flac/aacare now encoded from lossless PCM instead of Ogg Vorbis, and the defaultpcmoutput is resampled to 24 kHz for OpenAI parity (pass an explicitSampleRateto keep Polly's native rate). Same formats, different — better — bytes. - Error responses no longer expose backend internals. Server-side (
5xx) error messages are generic with details kept in the server log, and Anthropic error types now match the official SDK exactly. - Comprehend-backed moderation always analyses text as English — the only language the AWS API accepts at runtime.
- A configured guardrail now applies to every route. Embeddings, rerank, images, videos, and the audio routes enforce it through the ApplyGuardrail API (route coverage) — requests that silently bypassed the guardrail on v1.14 may now return
400(codecontent_filter) or masked text, and each check is billed as guardrail text units. - SSRF protection covers every non-globally-reachable address. With
SSRF_PROTECTION_BLOCK_PRIVATE_NETWORKSenabled (the default), a user-supplied URL resolving to shared address space (100.64.0.0/10, used by EKS custom networking and Hybrid Nodes) or another special-purpose range is now rejected with403, alongside the RFC 1918 ranges. - Usage reporting is additive but richer: cached-token buckets are folded into
prompt_tokenswithprompt_tokens_detailson every surface, and the Anthropic API now reportscache_creation_input_tokens(it was alwaysnullon v1.14). - Two request-body keys are reserved.
model_idandadditional_request_fields(plusstop_sequenceson the legacy/v1/completions, wherestopis the parameter to use) collide with the gateway's own request-building parameters: instead of being forwarded to Bedrock as provider extras, they return a400 invalid_request_errornaming the key. - The container health probe now respects
TRUSTED_HOSTS. The image'sHEALTHCHECKrequests/healthwith aHostheader derived fromTRUSTED_HOSTS— a correct list keeps the container healthy with no extra entry. Deployments that re-declare the probe, such as an ECS task definition, should run the image's own command. Note that a load balancer health check still sends the target's IP address as theHostand is rejected with400when the allow-list is enabled.
Explicit Prompt Caching¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
Explicit cache breakpoints on chat completions and responses, mapped to Bedrock cachePoint blocks with the per-request block budget enforced | ||
prompt_cache_options – cache TTL control on chat completions and responses, honoured on models that support it |
Prompt caching itself is not new — this release adds explicit placement and lifetime control on the OpenAI dialect, and fixes the caching plumbing that already existed (see Fixes).
Reasoning Controls¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| stdapi.ai | CHAT_COMPLETIONS_REASONING_FIELD – return thinking text under reasoning_content, reasoning, or suppress it with none; applied identically to streamed deltas and final messages | |
OpenRouter-style reasoning request object accepted on chat completions (effort, max_tokens, enabled, exclude) |
New API Features¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
Responses API prompt parameter – serve prompts stored in Bedrock Prompt Management, with versions and variables (opt-in, AWS_BEDROCK_ALLOW_PROMPT_ARN) | ||
| Mid-conversation system messages forwarded natively on Claude 4.8+ and Claude 5 family models instead of being folded into the system prompt | ||
| stdapi.ai | Guardrail asynchronous stream processing via the X-Amzn-Bedrock-GuardrailStreamProcessingMode request header | |
| stdapi.ai | Configured guardrails enforced on every route: embeddings, rerank, images, videos, and audio now apply them via the ApplyGuardrail API (route coverage) | |
Moderations amazon.bedrock-runtime-guardrail-checks model – inline guardrail content filter checks with no guardrail resource required, the new default fallback for omni-moderation-* in supported regions |
Speech & Audio¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/audio/speech – Polly SpeechMarkTypes for word/sentence/viseme/SSML timing marks | ||
/v1/audio/transcriptions – Amazon Transcribe extra parameters: multi-language identification, custom vocabularies, PII redaction, and more | ||
/v1/audio/transcriptions – gpt-transcribe context inputs: keywords and multi-language languages, with detected languages reported in the response | ||
/v1/audio/transcriptions & translations – any Converse-capable speech-input Bedrock model transcribes through a generic default (Voxtral rebuilt on the Converse API); uploads outside the accepted formats are transcoded automatically | ||
/v1/audio/translations – AWS Translate Formality, Profanity, Brevity, and custom terminologies |
Cohere Embed & Rerank¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
embedding_types – quantized (int8/uint8/binary/ubinary) and base64 embeddings on both embed routes; image embedding metadata reported | ||
Rerank v1 – structured JSON documents with rank_fields selection |
Platform Features¶
| Feature | Description |
|---|---|
| Stateless MCP transport | The MCP server can serve /mcp without server-side sessions (MCP_STATELESS_HTTP), so any replica may answer any request, alongside a GET /ping health probe kept out of request logs |
Retry-After on 429 | Throttled responses advertise the region router's computed backoff, so well-behaved clients retry exactly when capacity returns |
| AWS request-ID correlation | Request logs record every AWS API call's request ID (and incoming ALB/CloudFront trace headers), so a gateway request ties directly to CloudTrail and AWS support cases |
| Programmatic tool calling types | The OpenAI SDK's programmatic tool calling type surface parses on every request union, accepted and ignored on models without the capability |
| Performance | Hot paths run in compiled native code and independent work runs in parallel: a 1 MB request costs 30% less CPU, multi-image generations finish in the time of the slowest image, and every optimization is pinned by regression tests |
| MCP context efficiency | Tool schemas hide parameters MCP callers cannot use (streaming modes, token-level tuning, caller identifiers) and tool results return as compact JSON, cutting the tokens each call costs the calling agent; every exposed MCP tool is exercised end to end through a real MCP client in the test suite |
| Slimmer container image | The community image shrinks from 230 MB to 156 MB: unused dependency payloads are removed (language-name data, cryptography, rich/typer, uvicorn extras) and AWS service models are pruned to the services actually used, guarded by a build-time smoke test; the package inventory stays complete for vulnerability scanners (self-built ffmpeg registered, Python package metadata retained) and every redistributed component keeps its licence and notice files |
| Metadata filter for MCP clients | Listing stored chat completions accepts the metadata filter as a single metadata={"key": "value"} JSON object as well as the OpenAI SDK's metadata[key]=value pairs, so clients that can only send one query parameter per field — MCP tool calls among them — can filter too |
Verified with Real Clients¶
Compatibility claims in this release are backed by a new test tier that runs complete, unmodified third-party client software against a live gateway — not just HTTP assertions. Coding agents (Claude Code, Codex, pi, Qwen Code), the n8n workflow platform, Open WebUI, Home Assistant's voice bridge, a Haystack RAG pipeline, and the LangChain and pydantic-ai libraries drive real multi-turn tool-calling, retrieval, and speech sessions across dozens of models and all three API dialects, in isolated sandboxes. Alongside them, every served model is empirically probed for the parameters it genuinely honours, with the results recorded and pinned by tests. See Quality Assurance for the full methodology.
Fixes¶
Three audit passes and an independent full-branch review closed over a hundred fidelity gaps. The user-visible highlights:
- MCP tool calls match their schemas: union-typed parameters no longer advertise a contradictory single
type(which made valid string arguments randomly fail schema validation), and the JSON image edit/variation bodies accept the plain string references the tool schemas advertise - Clean errors on log-exempt paths: a 404 or 405 on paths kept out of request logs (e.g.
/favicon.ico, auto-requested by every browser visiting/) returned a 500 with a traceback instead of the JSON error envelope - Prompt caching plumbing: Anthropic
cache_controlbreakpoints land on their marked block; cache reads and writes are counted, priced, and reported consistently in responses, request logs, andcount_tokens - Reasoning:
reasoning_effort="max"accepted end to end; thinking text returned by Bedrock Mantle models is surfaced on every API instead of dropped; unsigned reasoning is no longer replayed to models that would reject it - Streaming parity: streamed and non-streamed results are now identical (tool-call indices, same-role message merging, text concatenation); mid-stream errors emit proper error events on every API instead of ending streams silently; redacted thinking and web-search results round-trip exactly as native Anthropic emits them; failed generations return
502instead of an empty200 - Routing & billing: a read timeout on an already-sent request is no longer re-invoked in another region (no double billing), on Converse and Mantle alike; region failover now tries each candidate region at most once per request instead of looping back over regions it just marked as throttled, so a single request can no longer escalate a region's quota backoff toward the one-hour ceiling; Bedrock prompt-router usage is billed against the actually-invoked model; the price card reprices a standard-tier row served as a tier fallback at the rate that tier actually bills;
store=truedegrades gracefully with a logged warning in regions without the session API - Responses API parity: the type surface is synchronized with the current OpenAI SDK (tool fields, error codes, tool-call
callerprovenance); Anthropiccount_tokenscounts exactly what generation sends, and error bodies carry therequest_id - Audio & images: the transcoding pipeline is fully bounded — a stalled or failed encode returns a clean error instead of holding the connection open; multipart forms bind every list field the OpenAI SDK sends;
size="auto"works on generation, edits, and variations;zh-TW/pt-PTstay distinct in translation; PII-redacted transcripts are read from the key Amazon Transcribe actually writes; Polly voice auto-selection is deterministic; batch-purpose files apply the documented 30-day default expiry, and an expired file now disappears from file listings instead of being listed with an entry that 404s on retrieve
v1.14.0 – 2026-07-12 – Bedrock Mantle, Video Generation, Cohere APIs, Moderation & Stored Conversations¶
At a glance
- Amazon Bedrock Mantle, enabled by default — the models served by the Mantle endpoint (OpenAI GPT-5.4/5.5/5.6, xAI Grok 4.3, Google Gemma 4, Qwen3, GLM, DeepSeek, MiniMax, Kimi, Nemotron and more) become available through all four text APIs, with independent throughput quotas.
- A Cohere-compatible API — Rerank and Embed, making this a three-dialect gateway.
- Videos and moderation — the OpenAI-compatible Videos API for asynchronous video generation, and content moderation backed by Amazon Bedrock Guardrails or Amazon Comprehend toxicity detection.
- Stored conversations —
store=true,previous_response_idcontinuation and a full lifecycle on Amazon Bedrock session storage, plus conversation compaction and Responses API extended reasoning. - Operations — a model pricing API, multi-region failover for every AWS AI service, fault-tolerant startup, real AWS-billed usage and costs in request logs, and a security hardening pass. Two new IAM permissions are required.
This release adds enabled-by-default Amazon Bedrock Mantle support — models served by the Bedrock Mantle endpoint (OpenAI GPT-5.4/5.5/5.6, xAI Grok 4.3, Google Gemma 4, Qwen3, GLM, DeepSeek, MiniMax, Kimi, Nemotron, and more) become available through all four text APIs, with transparent API conversion, native stored conversations, and independent throughput quotas. It also turns stdapi.ai into a three-dialect gateway with the new Cohere-compatible API (Rerank and Embed), adds the OpenAI-compatible Videos API for asynchronous video generation, content moderation backed by Amazon Bedrock Guardrails or Amazon Comprehend toxicity detection, stored responses and chat completions with store=true, previous_response_id multi-turn continuation, and a full list/retrieve/update/delete lifecycle on Amazon Bedrock session storage, and conversation compaction. The Responses API gains extended reasoning: Bedrock reasoningContent now surfaces as native reasoning output items, both non-streaming and streamed, with signatures and redacted payloads round-tripping through an encrypted_content envelope. A broader compatibility pass brings request/response parity closer to the OpenAI SDK — hosted and agent tool types (web search, computer use, custom tools) are now accepted and ignored instead of rejected, streams correctly terminate with response.incomplete/response.failed, cached tokens are counted in input_tokens, and citation annotations are emitted with their streaming events — validated end-to-end against the OpenAI Codex CLI as an agent client. Operations gain a model pricing API, multi-region failover for every AWS AI service, fault-tolerant startup, real AWS-billed usage and costs in request logs (optionally exported as CloudWatch metrics), and a security hardening pass covering SSRF protection, input validation, and log/error redaction.
New Required IAM Permissions
v1.14.0 requires two new IAM permissions:
bedrock:Rerank— needed for the Cohere-compatible Rerank API (/cohere/v2/rerank). See IAM Permissions.bedrock:ListAsyncInvokes, plusbedrock:ListTagsForResourceonarn:aws:bedrock:*:*:async-invoke/*— needed forGET /v1/videos(listing video generation jobs across regions). See IAM Permissions.
Ensure your IAM role or user policy includes both statements before upgrading to v1.14.0.
Session storage and Comprehend permissions already covered
The IAM permissions for stored responses/chat completions (bedrock:CreateSession and related session actions) and Comprehend-based moderation (comprehend:DetectToxicContent) were already added to the official stdapi-ai Terraform module ahead of this release. Deployments using a hand-written policy still need to add those statements if they haven't already. Without the session permissions, store=true (previously accepted and ignored) is still ignored — a warning is recorded in the request log instead of failing the request.
Amazon Bedrock Mantle¶
Enabled-by-default support (AWS_BEDROCK_MANTLE_ENABLED) for models served by the Amazon Bedrock Mantle endpoint — OpenAI GPT-5.4/5.5/5.6 (Sol, Terra, Luna), xAI Grok 4.3, Google Gemma 4, Qwen3, GLM 4.x/5, DeepSeek V3.x, MiniMax M2.x, Kimi K2.5, Nemotron, and more — alongside the classic Bedrock Converse catalog:
- All four text APIs (chat completions, responses, messages, legacy completions) are served for every Mantle model — native passthrough where the model supports the API upstream, transparent conversion otherwise
- Models available on both bedrock-runtime and Mantle are served by bedrock-runtime by default;
AWS_BEDROCK_MANTLE_PREFERRED_MODELSor the opt-inx-stdapi-servicerequest header (AWS_BEDROCK_MANTLE_SERVICE_HEADER) route them through Mantle instead — e.g. to tap Mantle's independent throughput quotas - Native Mantle stored conversations on
/v1/responses(store,previous_response_id, retrieval and deletion) — 30-day retention, region-local, project-scoped - Multi-region failover and quota backoff across
AWS_BEDROCK_MANTLE_REGIONS, matching classic Bedrock region routing - Authentication via short-term bearer tokens derived from the server's AWS credential chain — no static secrets
- Usage recorded and priced at bedrock-mantle rates, including cached tokens and service tiers
- Optional Bedrock Project/Workspace attribution for cost tracking via
AWS_BEDROCK_MANTLE_PROJECT, with per-request override (AWS_BEDROCK_ALLOW_MANTLE_PROJECT_OVERRIDE) through theOpenAI-Project/anthropic-workspaceheader
Additional IAM Permissions (opt-in feature)
Enabling AWS_BEDROCK_MANTLE_ENABLED requires the bedrock-mantle:CreateInference, bedrock-mantle:GetInference, bedrock-mantle:DeleteInference, bedrock-mantle:ListModels, bedrock-mantle:GetModel, and bedrock-mantle:CancelInference permissions on arn:aws:bedrock-mantle:*:*:project/*, plus bedrock-mantle:CallWithBearerToken on *. See IAM Permissions.
New APIs¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/videos – create, poll, list, download, and delete video generation jobs | ||
/v1/moderations – text and image content classification | ||
/cohere/v2/rerank – document reranking (Amazon Rerank 1.0, Cohere Rerank 3.5) | ||
/cohere/v2/embed – embeddings over all Bedrock embedding models | ||
| stdapi.ai | /model_pricing – exact AWS unit prices per model | AWS Price List API |
Extended Reasoning¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/responses – Bedrock reasoningContent returned as reasoning output items | ||
Streaming response.output_item.added / response.reasoning_text.delta / .done events for reasoning content | ||
include=["reasoning.encrypted_content"] – signature/redacted round-trip for multi-turn reasoning continuation |
Conversations¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
store=true + GET/DELETE /v1/responses/{id}, input items listing, and previous_response_id continuation | ||
POST /v1/responses/{id}/cancel – endpoint parity for the cancel lifecycle (always fails for session-stored responses, which never run in background mode; Mantle-stored responses are cancelled upstream) | ||
store=true + GET/DELETE /v1/chat/completions/{id}, GET /v1/chat/completions listing, POST /v1/chat/completions/{id} metadata updates, and input messages listing | ||
/v1/responses/compact – stateless conversation compaction | ||
moderation request parameter on chat completions and responses, with results reported in the response |
Platform Features¶
| Feature | Description |
|---|---|
| Claude 5 models | Explicit support for the Claude 5 generation — Opus 5, Sonnet 5, Fable 5, and Mythos — with each model's server tool set and reasoning configuration matched to what Bedrock actually accepts (Opus 5 exposes no computer use tool; Fable and Mythos always reason and reject a disabled configuration). Model matching covers unreleased versions of each family, so a new minor or major release inherits its family's behavior instead of a generic fallback. Validated end-to-end across the full Claude feature matrix, from Claude 4.5 through Claude 5 |
| Multi-region AWS AI services | Automatic multi-region failover for Amazon Polly, Transcribe, Translate, and Comprehend (per-engine voice discovery, co-located Transcribe buckets, latency-ordered region pools) |
| Fault-tolerant, faster startup | Unreachable Bedrock regions or Polly engines no longer abort startup; they are skipped with a warning and retried on the next refresh — and startup is faster overall |
| Usage & cost tracking | Request logs report the usage actually billed by AWS with its cost computed from live AWS pricing, optionally exported as CloudWatch metrics (CLOUDWATCH_METRICS); the previous token estimation is removed and its TOKENS_ESTIMATION* settings are deprecated and ignored |
| Cost attribution | Request, server, and user correlation metadata is now attached to every synchronous Bedrock inference call — the InvokeModel family included, not only Converse — so Bedrock invocation logs can be filtered and costs attributed per request or per user |
| Smaller container images | The published images shrink by around 40% — 413 MB to 253 MB for the AWS Marketplace image, 377 MB to 230 MB for the community image — cutting pull time and storage. ffmpeg is now built with only the audio encoders the server uses, and the unused OpenTelemetry gRPC exporter is no longer installed |
Video retention (AWS_S3_VIDEOS_EXPIRES_AFTER) | Optional retention period for generated videos, reported as expires_at and enforced on download |
Upload expiry (expires_after) | Multipart upload sessions honor the OpenAI expires_after policy on the resulting file |
| Session storage encryption | Optional KMS key for Amazon Bedrock session storage (AWS_BEDROCK_SESSION_ENCRYPTION_KEY_ARN) |
Proxy trust (PROXY_TRUSTED_HOSTS) | X-Forwarded-* headers are only honored when sent by a trusted reverse-proxy address |
Input file size limit (MAX_INPUT_FILE_SIZE) | Optional cap on the size of downloaded/decoded input files, with bounded download concurrency (MAX_CONCURRENT_INPUT_DOWNLOADS) |
| Legacy model opt-in fix | AWS_BEDROCK_LEGACY now also exposes models whose AWS legacy date has already passed (e.g. Amazon Nova Reel) |
Security Hardening¶
- MCP transports and
/search_modelsnow require authentication when an API key is configured — clients that relied on these endpoints being open must now send the API key - SSRF protection hardened against IP-literal encoding and DNS-rebinding bypasses on URL file inputs
s3://file inputs are restricted to the server's allowed buckets, and multipart upload filenames are validated- Decoded image size is capped against decompression-bomb payloads
- ARNs and AWS account IDs are redacted from client-facing error messages, and presigned URL signatures are stripped from logs and traces
- An empty resolved API key (e.g. a blank secret value) now disables authentication cleanly instead of matching an empty bearer token, and CORS no longer allows credentialed cross-origin requests
- Reduced container attack surface: the images no longer carry the video codec, X11 and font libraries that a distribution ffmpeg package links — x264, x265, AOM, dav1d, SVT-AV1 and SDL2 among them — nor the gRPC stack. None were reachable from the audio transcoding ffmpeg is used for, and they accounted for the bulk of the images' third-party native code. ffmpeg is built from the same version the base distribution ships, with only the audio encoders in use, and enables no GPL-licensed component
Agent SDK Compatibility¶
The Responses API request/response surface was audited and hardened against the OpenAI SDK and real agent clients, end-to-end tested against the OpenAI Codex CLI:
- Hosted and agent tool types (
web_search,computer_use,file_search,custom/namespacetools, and other items without a Bedrock equivalent) are now accepted and dropped instead of rejected with400, preserving compatibility with existing agent tooling - Streaming responses now correctly terminate with
response.incompleteorresponse.failed(matching upstream behavior) instead of always reportingresponse.completed - Mid-stream errors emit the spec-compliant
errorSSE event input_tokensusage now includes cache read/write tokens, matching OpenAI's accountingurl_citationannotations are emitted alongside their streaming events- Echoed reasoning items tolerate the field variations produced by different SDKs and agent clients
Fixes¶
- Rerank models are no longer incorrectly advertised on Converse-based chat routes and MCP tools
- Fixed per-request model parameter overrides (
default_model_params) occasionally leaking into subsequent requests for the same model - 5xx provider errors now report server-side error types (
server_error/api_error) in OpenAI and Anthropic error envelopes instead ofinvalid_request_error - Unknown paths (
404) and wrong methods (405) now return the error envelope of the API family they were sent to, instead of the framework's defaultdetailpayload - The Anthropic Messages API now returns
404instead of400for an unknown model, matching the upstream API, and rejects atop_pabove1.0 - Audio transcription returns plain text for
response_format=textand now defaultsverbose_jsonto segment timestamps - Responses API usage reports
input_tokens_details.cache_write_tokens, which recent OpenAI SDKs require to parse a response - Files API listing and cursor pagination order by creation time again: file IDs now use an order-preserving alphabet, where the previous one could sort a newer file first. IDs issued before this release keep working, but sort among themselves as before until they expire
- Newer Anthropic client request fields (free-form JSON Schema keywords in tool
input_schema, adaptive thinkingdisplay) are accepted instead of rejected in strict validation mode - Amazon Nova 2 no longer fails on
max_tokenscombined with high reasoning effort (the cap is dropped with a logged warning) - Explicit cache points are kept off tool-related content blocks for models without tool caching support
- The Files API unavailable error no longer exposes the S3 bucket configuration detail
- Fixed input files from one request occasionally leaking into later requests served by the same connection, which could fail those requests with internal errors
- Anthropic Messages streams now emit an empty tool-input delta for tool calls without arguments, so SDK stream accumulators no longer fail on argument-less tool calls
- JSON-body image edit and variation requests now accept the
modelfield instead of rejecting the request - Model listings now report
service: "AWS Bedrock Runtime"for classic Bedrock models (previously"Amazon Bedrock"), distinguishing them from"AWS Bedrock Mantle" - High reasoning effort now maps to the intended thinking-token budget on Anthropic Claude models (the budget factor was previously miscomputed)
- Setting
log_leveltodisablednow suppresses all log output as documented, instead of publishing every event - Server startup no longer fails when the ECS container metadata endpoint answers slowly, which could prevent small Fargate tasks from starting: the lookup is retried, then falls back to the STS caller identity with a startup warning
- Multipart upload parts are numbered from the parts already stored in S3 instead of a per-instance counter: with several server instances behind a load balancer, two parts of one upload could be given the same number, overwriting each other and failing the upload
- Multi-region failover now covers a region that does not offer the service at all: with no
AWS_COMPREHEND_REGIONset, a Bedrock region without Amazon Comprehend moves on to the next one as documented, instead of failing language detection and Comprehend moderation
v1.13.0 – 2026-07-03 – Terraform Module Compliance & Security Hardening¶
At a glance
- A Terraform-module release — the stdapi-ai module and its VPC, KMS and ECS Fargate children, with no server change.
- Security Hub FSBP control documentation — every module README now carries a full Foundational Security Best Practices control mapping.
- Compliance gaps closed — default security group lockdown, ALB access logging, and EFS POSIX user enforcement with native backups.
- Optional network integrations — compliance VPC endpoints, a GuardDuty VPC endpoint, Route 53 Resolver DNS Firewall, and VPC Flow Logs retention.
- Tagging and token cost — all four modules accept a
tagsvariable, and MCP tool descriptions were shrunk, lowering the token cost of every agent session.
This release focuses on the stdapi-ai Terraform module and its child modules — VPC, KMS, and ECS Fargate — adding detailed AWS Security Hub control documentation and closing several compliance gaps: default security group lockdown, ALB access logging, EFS POSIX user enforcement with native backups, and optional compliance/GuardDuty/DNS Firewall VPC integrations. All four modules now also accept a tags variable for custom resource tagging.
Documentation-first release
Every module README now includes a full Security Hub Foundational Security Best Practices (FSBP) control mapping. See Authentication & Security for a summary and links to each module.
Fixes¶
- Added the missing
1hand5mvalues toPromptCacheRetentionfor Bedrock-specific prompt cache TTLs in the OpenAI Responses API
Security Hub & Compliance Hardening¶
| Feature | Module | Description |
|---|---|---|
| Security Hub FSBP control documentation | VPC, KMS, ECS Fargate, stdapi-ai | Per-control (pass/fail/conditional/N-A) tables added to each module README |
| Default security group lockdown | VPC | New aws_default_security_group resource revokes all default ingress/egress rules (EC2.2 / CIS 5.4) |
| VPC Flow Logs retention | VPC | Default retention increased from 7 to 365 days (EC2.6) |
| Compliance VPC endpoints | VPC | New compliance_vpc_endpoints_enabled variable adds ECR, SSM, SSM Contacts, and SSM Incidents interface endpoints |
| GuardDuty VPC endpoint | VPC | New guardduty_vpc_endpoint_enabled variable adds the guardduty-data interface endpoint |
| Route 53 Resolver DNS Firewall | VPC | New dns_firewall_enabled variable blocks/alerts on DNS queries to known-malicious domains (AWS Managed Domain Lists, plus DGA/DNS-tunneling detection via dns_firewall_advanced_enabled); dedicated VPC only |
| ALB access logging | stdapi-ai | New alb_access_logging_enabled variable (default true) logs ALB access to a dedicated, encrypted S3 bucket |
| EFS POSIX user enforcement | ECS Fargate | mount_points now accepts an efs_posix_user object to enforce a POSIX identity on EFS access points (EFS.4) |
| EFS native backups | ECS Fargate | New mount_points_efs_backup_enable variable enables native EFS automatic backups, independent of the existing AWS Backup plan (EFS.7) |
| Resource tagging | VPC, KMS, ECS Fargate, stdapi-ai | New tags variable propagates custom tags to nearly all created resources (IAM.24 / EC2.48) |
Other Infrastructure Changes¶
| Feature | Description |
|---|---|
| AWS provider version bump | Requirement raised to >= 6.27.0 across all four modules |
| S3 object tag rename | Files API objects and the corresponding Terraform lifecycle rule now use the stdapi-ai.expires tag key instead of expires; a temporary backward-compatible rule still expires legacy-tagged objects |
aws-apn-id resource tagging | AWS resources created at runtime (Bedrock async jobs, Transcribe jobs, S3 objects) are tagged with aws-apn-id, the standard AWS Marketplace attribution tag — an internal, vendor-side tag, not user-configurable |
MCP Token Optimization¶
- Significantly reduced the size of MCP tool descriptions across the API, lowering the token cost of every AI agent session connected to this server
- No change in functionality: all parameter constraints and usage guidance remain intact
v1.12.0 – 2026-05-29 – Completions API, Video Understanding & File References¶
At a glance
/v1/completions— the OpenAI text completion endpoint, for text-first coding agents and legacy completion clients.- Video understanding — TwelveLabs Pegasus analyses
video/*inputs in chat completions, honouringservice_tierand guardrail configuration. - Input token counting —
/v1/responses/input_tokensfor the Responses API. - The
file-id:URI scheme — reference a Files API upload anywhere a URL is accepted: embeddings, transcription, chat, images and messages. - Settings and compatibility —
DEFAULT_MODEL_SERVICE_TIERSapplies a per-model service tier automatically, reasoning can be explicitly enabled or disabled, and the Anthropic/v1/messagesroute acceptssystem-role messages.
This release adds the OpenAI-compatible /v1/completions endpoint for text-first coding agents and legacy completion clients, TwelveLabs Pegasus video understanding for analyzing video/* inputs in chat completions, and an input token counting endpoint for the Responses API. Files uploaded through the Files API can now be referenced anywhere a URL is accepted using the new file-id: URI scheme. The Anthropic Messages API now accepts system-role messages (merged into the system prompt for compatibility), reasoning can be explicitly enabled or disabled, and a new DEFAULT_MODEL_SERVICE_TIERS setting applies per-model service tiers automatically.
Chat Completions¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/completions – text completion endpoint for text-first coding agents | ||
/v1/responses/input_tokens – input token counting | ||
/v1/messages – accepts system-role messages (merged into the system prompt) | ||
Pegasus video understanding (video/* inputs) |
Speech & Audio¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/audio/speech – case-insensitive voice names & default model |
Platform Features¶
| Feature | Description |
|---|---|
file-id: URI scheme | Reference Files API uploads via file-id:<file-id> anywhere a URL is accepted — embeddings, audio transcription/translation, chat, images, and messages |
Default model service tiers (DEFAULT_MODEL_SERVICE_TIERS) | Automatically apply a per-model service tier (default, flex, priority, reserved) when none is provided in the request |
| Explicit reasoning enable/disable | Reasoning/thinking can now be explicitly enabled or disabled via request parameters |
| Service tier & guardrail support for Pegasus | TwelveLabs Pegasus requests honor service_tier and Bedrock Guardrail configuration |
| MCP speech streaming defaults to SSE | /v1/audio/speech defaults stream_format to sse when invoked as an MCP tool for broader client compatibility |
| Full regional S3 bucket handling | The Terraform module resolves regional S3 buckets via resource-level region (requires AWS provider >= 6.0.0) |
| Reliable cross-region model identifiers | Region routing no longer fails intermittently with "The provided model identifier is invalid": a region whose inference profile is missing or not yet propagated is skipped, and a geo-scoped profile is never sent to a different region |
v1.11.0 – 2026-05-02 – MCP Server, Agent Discovery & Model Search (with v1.11.1–v1.11.4 maintenance updates, through 2026-05-28)¶
At a glance
- An MCP server — every API endpoint is exposed as an MCP tool, over Streamable HTTP and SSE transports that are independently enabled and selectively restricted.
/search_models— filter models by route, MCP tool name, input and output modalities, region, streaming support and legacy status.- Agent discovery — RFC 8288 Link headers on
/, an RFC 9727 API catalog at/.well-known/api-catalog, an MCP Server Card, androbots.txtcontent signals. - JSON bodies for binary endpoints — audio transcription, audio translation and image edits accept
application/jsonwith files as base64, data URI, HTTP URL or S3 URI. - Maintenance (v1.11.1–v1.11.4) —
max_tokensmade optional on Anthropic/v1/messages, MCP dependencies added to the container image, and Starlette upgraded for CVE-2026-48710.
This release introduces a Model Context Protocol (MCP) server, making all stdapi.ai API endpoints directly accessible as MCP tools for AI agents and agentic workflows. A new /search_models endpoint enables precise discovery of models by route, MCP tool, region, streaming support, and legacy status. Agent-friendly discovery metadata is now exposed via RFC 8288 Link headers and an RFC 9727 machine-readable API catalog at /.well-known/api-catalog. Endpoints that previously required binary multipart/form-data uploads now also accept an application/json body for MCP and HTTP client compatibility. The Anthropic Messages API now accepts xhigh as a reasoning_effort value.
MCP Server¶
| Feature | Description |
|---|---|
| MCP server (Streamable HTTP & SSE) | All API endpoints exposed as MCP tools; Streamable HTTP and SSE transports can be independently enabled or disabled via configuration |
| Configurable MCP tool exposure | Individual MCP tools can be selectively enabled or restricted via configuration |
| JSON body for binary endpoints | Audio transcription, audio translation, and image edit endpoints now accept application/json with files as base64, data URI, HTTP URL, or S3 URI |
Model Search¶
| Feature | Description |
|---|---|
/search_models | New official endpoint to filter models by route, MCP tool name, input/output modalities, region, streaming, and legacy status; returns richer metadata than /v1/models or Anthropic /v1/models, designed for LLM-driven model selection (replaces BETA and undocumented /available_models) |
Agent Discovery¶
| Feature | Description |
|---|---|
| RFC 8288 Link headers | Root (/) endpoint returns Link headers for resource discovery |
RFC 9727 API catalog (/.well-known/api-catalog) | Machine-readable API catalog for automated agent and tool discovery |
MCP Server Card (/.well-known/mcp/server-card.json) | Advertises available MCP transports and capabilities to AI agents (SEP-1649) |
robots.txt AI signals | Updated robots.txt with Content-Signal directives and explicit /.well-known/ allow rule |
Chat Completions & Messages¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/messages reasoning_effort=xhigh support |
Deprecation Mappings¶
- Added automatic fallback for
amazon.nova-reel-v1:0andanthropic.claude-3-haiku-20240307-v1:0to their respective replacements
Fixes¶
- Fix reasoning token double-counting in usage calculation in OpenAI Responses API adapter
- Fix missing
file_idinputs for image and file processing in OpenAI Responses API adapter - Remove
storeparameter from unsupported validations in chat completions to ensure client compatibility
Fixes & Maintenance (v1.11.1–v1.11.4)¶
v1.11.1
- Make
max_tokensoptional in Anthropic/v1/messagesto align with the Anthropic API specification - Remove unsupported reasoning configuration checks for broader client compatibility
- Rename
/v1/responsesroute tag from "Responses" to "Chat" in OpenAPI documentation for consistency
v1.11.2-v1.11.3
- Add missing MCP dependencies to container image.
v1.11.4
- Upgrade Starlette dependency to fix CVE-2026-48710.
v1.10.0 – 2026-04-17 – OpenAI Responses API¶
At a glance
/v1/responses— OpenAI's API for agents and multi-step workflows, drop-in compatible with the OpenAI SDK.- Every Converse-compatible model — it works with all Amazon Bedrock Converse-compatible models, streaming included.
- Built-in tools —
web_search/web_search_preview,code_interpreterandimage_generation. - Function tools, extended reasoning and structured output on the same surface.
- Fixes — prompt caching with tool-related content, an optional
signaturefield in Anthropic message types, and model legacy detection when the end-of-life date falls before the next cache refresh.
This release adds support for the OpenAI /v1/responses endpoint—OpenAI's next-generation API designed for building agents and multi-step AI workflows. Drop-in compatible with the OpenAI SDK, it works with all Amazon Bedrock Converse-compatible models and supports streaming, function tools, built-in tools (web search, code interpreter, image generation), extended reasoning, and structured output.
Responses (OpenAI-Compatible)¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/responses | ||
/v1/responses – web_search / web_search_preview built-in tool | ||
/v1/responses – code_interpreter built-in tool | ||
/v1/responses – image_generation built-in tool |
Fixes¶
- Fix prompt caching error when messages contain tool-related content on models that do not support tool caching
- Make
signaturefield optional in Anthropic message types - Fix model legacy detection when the end-of-life date falls before the next cache refresh
v1.9.0 – 2026-04-10 – Files API & Images API JSON Body¶
At a glance
- A Files API backed by Amazon S3 —
/v1/filesCRUD on both the OpenAI-compatible and Anthropic-compatible interfaces, sharing one store. - Incremental uploads —
/v1/uploads, the OpenAI multipart uploads API, for large files. - File IDs as model inputs — a stored file is usable as a document or image input in chat completions and in messages.
- A JSON body for image editing —
/v1/images/editsand/v1/images/variationsacceptapplication/jsonreferencing Files API IDs or URLs, so pipeline steps chain without re-uploading. - New required configuration —
AWS_S3_BUCKETmust be set, with read, write, delete and list permissions on it.
This release introduces a Files API backed by Amazon S3, available through both the OpenAI-compatible and Anthropic-compatible interfaces. Files uploaded via either API share the same S3 storage and can be referenced across both interfaces. Large files can be uploaded incrementally using the OpenAI multipart uploads API. Stored files can be referenced by ID directly in image edit and variation requests (JSON body), as well as in chat completion messages as document or image inputs. The image edits endpoint now also accepts an application/json body as an alternative to multipart form-data, making it easier to chain pipeline steps without re-uploading files.
New Required Configuration
Files API requires AWS_S3_BUCKET to be configured (shared with the image URL response feature). The S3 prefix for stored files defaults to files/ and is configurable via AWS_S3_FILES_PREFIX. Ensure your IAM role includes read, write, delete, and list permissions on the files prefix in addition to the existing S3 permissions for presigned URLs.
Files & Storage¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/files – CRUD operations | ||
/v1/uploads – multipart uploads | ||
/v1/files – CRUD operations |
Image Generation¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/images/edits – JSON body with images/mask referencing Files API IDs or URLs | ||
/v1/images/variations – JSON body with image referencing a Files API ID or URL |
Chat Completions & Messages¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| Files API file IDs usable as document/image inputs in chat completions | ||
| Files API file IDs usable as document/image inputs in messages |
Fixes¶
- Document inputs via S3 URLs are not supported as Bedrock Converse API inputs for some models (e.g., Claude) — now properly detected and handled
v1.8.0 – 2026-04-04 – Broader Model Compatibility & Structured Output¶
At a glance
- Structured output —
response_formatwith JSON object and JSON schema on OpenAI chat completions. - Request metadata —
metadatais forwarded to Bedrock, and the request context (request_id,server_id,user_id) is tagged onto every Bedrock and Amazon Transcribe job. - Tool handling — Amazon Nova's grounding tool maps to
web_searchcontent blocks, with multi-turn support, and the brokensystemTool_auto-promotion was removed. - Region routing — region-restricted models always get non-global inference profiles, and the case where no region is usable is handled gracefully.
- New required IAM permissions —
bedrock:TagResourceandtranscribe:TagResource;AWS_BEDROCK_LEGACYnow defaults tofalse.
This release focuses on improving reliability and compatibility across a wide variety of models. Structured response formats (JSON object and JSON schema) are now supported on OpenAI chat completions, and request metadata can be forwarded to Bedrock. Tool handling has been significantly improved—both for model-specific system tools and for Amazon Nova's grounding tool, including multi-turn support. Region routing is now more robust, correctly enforcing non-global inference profiles for region-restricted models and handling edge cases gracefully.
New Required IAM Permissions
v1.8.0 requires two new IAM permissions to attach request metadata tags to jobs:
bedrock:TagResourceonarn:aws:bedrock:*:*:async-invoke/*— needed for Bedrock asynchronous invocation jobs (see IAM Permissions). Thetwelvelabs.marengo-embed-3-0-v1:0andtwelvelabs.marengo-embed-2-7-v1:0models rely on asynchronous invocation and will fail with an access denied error if this permission is missing.transcribe:TagResourceonarn:aws:transcribe:*:*:transcription-job/*— needed for Amazon Transcribe transcription jobs (see IAM Permissions). Theamazon.transcribemodel will fail with an access denied error if this permission is missing.
Ensure your IAM role or user policy includes both statements before upgrading to v1.8.0.
Chat Completions¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
response_format – JSON object and JSON schema structured output | ||
metadata – request metadata forwarding to Bedrock | ||
| Nova Code Interpreter global profile support |
Messages (Anthropic-Compatible)¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
nova_grounding responses mapped to web_search content blocks | ||
Multi-turn conversation support with nova_grounding |
Platform Features¶
| Feature | Description |
|---|---|
| Non-global profiles for region-restricted models | Region-restricted models are now always assigned non-global inference profiles, preventing requests from bypassing configured region restrictions |
| Region routing edge case handling | Region routing gracefully handles cases where no usable regions are available |
| ECS-based server ID | When running on ECS, server_id in logs is set to task_id.container_name for precise instance identification across tasks and containers |
| Request metadata tagging | stdapi.ai request context (request_id, server_id, user_id) is automatically attached as tags to every Bedrock and Amazon Transcribe job, making it easy to trace API calls across AWS service logs |
Fixes¶
- Fix
systemTool_prefix handling: removed broken auto-promotion logic; system tools require specific tool output handling not compatible with generic tool forwarding AWS_BEDROCK_LEGACYdefault changed fromtruetofalseto prevent access denied errors on legacy models that have not been actively used recently- Bedrock read timeouts are now handled as standard model errors (503) instead of unhandled exceptions, and are properly retried across regions when multi-region routing is enabled
v1.7.0 – 2026-03-20 – Automatic Region Routing, Deprecated Model Fallback & Resilience Improvements¶
At a glance
- Automatic multi-region routing — Bedrock requests are distributed across the configured AWS regions, failing over on quota limits or unavailability.
- More quota by adding regions — each region carries its own quota, so adding one multiplies the effective tokens-per-minute and daily limits.
- Deprecated model fallback — deprecated model IDs are transparently rerouted to their replacements, with an extensible mapping, so clients survive AWS model retirements unchanged.
- A configurable AI response timeout, so a model call cannot hang indefinitely.
- S3 URLs for file inputs across all relevant endpoints, alongside HTTP URLs, data URIs and base64, with memory-efficiency improvements.
The headline feature of v1.7 is automatic multi-region routing: stdapi.ai now intelligently distributes requests across your configured AWS regions, failing over automatically on quota limits or unavailability—and because each region carries its own independent quota, adding regions directly multiplies your effective tokens-per-minute and daily limits. Alongside this, deprecated model IDs are transparently redirected to their replacements so clients survive AWS model retirements without any code changes. This release also adds S3 URL support for file inputs across all relevant endpoints, a configurable AI response timeout, and memory efficiency improvements.
Platform Features¶
| Feature | Description |
|---|---|
| Automatic region routing with configurable strategies | Intelligently distributes Bedrock requests across configured AWS regions with automatic failover on quota limits or unavailability; supports ordered, lowest_latency, and round_robin strategies |
| Deprecated model fallback | Transparently reroute deprecated model IDs to their replacements; extend or override the built-in mapping; warns on legacy model usage |
| AI response timeout | Configurable timeout for AI model responses to prevent indefinitely hanging requests |
| Expanded file input support | File inputs (images, documents, audio) now support S3 URLs in addition to HTTP URLs, data URIs, and plain base64 across all relevant endpoints; improves memory efficiency by releasing file data as early as possible |
| Model lifecycle timestamps | Model created/updated timestamps now derived from lifecycle data (startOfLifeTime, endOfLifeTime) |
Fixes¶
- Fix SSE stream error handling in monitoring to handle specific API and AWS client errors gracefully
- Fix audio MIME type detection failure when
libmagic's in-memory buffer path silently returnsapplication/octet-stream; fall back to file-based detection to ensure correct format is sent to Bedrock
v1.6.0 – 2026-02-27 – Anthropic API Compatibility & Advanced Claude Capabilities¶
At a glance
- A full Anthropic-compatible API —
/v1/messagesand/v1/messages/count_tokens, usable straight from the Anthropic SDK. - Anthropic-format model discovery —
/v1/modelsand/v1/models/{model_id}. - Claude server tools — bash, text editor, computer and memory, on both the Messages surface and OpenAI chat completions; Amazon Nova's
web_searchmaps tonova_grounding. - Configurable route prefixes —
ANTHROPIC_ROUTES_PREFIXandOPENAI_ROUTES_PREFIX, plus Anthropic beta flag filtering to prevent BedrockValidationExceptionerrors. - Real usage tracking — token counts sourced directly from AWS billing data instead of tiktoken estimation, and Claude model name aliases resolved to Bedrock identifiers.
Introduces a full Anthropic-compatible API layer, enabling direct use of the Anthropic SDK and Claude-native tools with Amazon Bedrock. Adds Claude server tools support via OpenAI chat completions, token count estimation, automatic Anthropic beta flag filtering, and configurable route prefixes.
Chat Completions¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/chat/completions Claude server tools (bash, str_replace_based_edit_tool, computer, memory) |
Messages (Anthropic-Compatible)¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/messages – Full Anthropic Messages API | ||
/v1/messages/count_tokens – Token counting | ||
| Claude server tools (bash, text editor, computer, memory) | ||
Web search tool (web_search → nova_grounding) |
Model Discovery (Anthropic-Compatible)¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/models – List models (Anthropic format) | ||
/v1/models/{model_id} – Get model details |
Platform Features¶
| Feature | Description |
|---|---|
ANTHROPIC_ROUTES_PREFIX configuration | Configurable base path prefix for Anthropic-compatible routes (default: /anthropic) |
OPENAI_ROUTES_PREFIX configuration | Configurable base path prefix for OpenAI-compatible routes |
Real usage tracking (usage in logs) | Token counts sourced directly from AWS billing data (replaces tiktoken-based estimation) |
Anthropic beta flag filtering (ANTHROPIC_BETA_FILTER) | Automatically filter unsupported anthropic-beta flags to prevent Bedrock ValidationException errors; extensible via ANTHROPIC_BETA_ALLOWLIST |
| Claude model name aliases | Use official Anthropic model names (e.g., claude-opus-4-8) auto-resolved to Amazon Bedrock identifiers |
v1.5.0 – 2026-02-15 – Advanced Reasoning & Model Compatibility (with v1.5.1–v1.5.2 maintenance updates, through 2026-02-18)¶
At a glance
- Amazon Nova 2 reasoning implemented on chat completions.
- Claude 4.6+ adaptive reasoning configuration.
- System prompt handling for models that do not support one, widening model compatibility.
- v1.5.1 — Amazon Nova Canvas image editing falls back to the
TEXT_IMAGEtask type when no mask is provided. - v1.5.2 — a
/route so the root endpoint stops answering404, and empty system content blocks are handled for Converse API compatibility.
Introduces advanced reasoning capabilities with Amazon Nova 2 and Anthropic Claude 4.6+ adaptive reasoning, enhanced system prompt handling for broader model compatibility.
Chat Completions¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| System prompt handling for unsupported models | ||
| Nova 2 chat model reasoning implementation | ||
| Claude 4.6+ adaptive reasoning configuration |
Fixes & Maintenance (v1.5.1–v1.5.2)¶
v1.5.2
- Add "/" route to avoid 404 errors on root endpoint
- Fix empty system content block handling (improves Amazon Bedrock Converse API compatibility)
v1.5.1
- Fix Amazon Nova Canvas image editing to fall back to TEXT_IMAGE task type when no mask is provided
v1.4.0 – 2026-02-11 – Audio Enhancements & Model Compatibility¶
At a glance
- Mistral Voxtral joins the audio models.
- Speaker diarization — the
diarized_jsonformat on/v1/audio/transcriptions. - Audio formats on chat completions, with extended Bedrock finish-reason mapping.
- Prompt caching TTL support on chat completions.
- Model aliasing — OpenAI-style model names resolved for seamless compatibility.
Expands audio capabilities with Mistral Voxtral support, speaker diarization, audio formats for chat completions, and introduces prompt caching TTL and model aliasing for better OpenAI compatibility.
Chat Completions¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/chat/completions audio format support | ||
/v1/chat/completions extended Bedrock finish reasons mapping | ||
| Prompt caching TTL support |
Speech & Audio¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/audio/transcriptions diarized_json format | ||
| Voxtral audio model |
Platform Features¶
| Feature | Description |
|---|---|
| Model alias support | Seamless OpenAI compatibility via model name aliasing |
Fixes¶
- Fix chat completion file input handling and refactor base64 decoding and MIME handling for file processing.
- Re-raise startup exceptions and disable botocore logging to improve error visibility
v1.3.0 – 2026-01-11 – Image Editing & Variation Support (with v1.3.1–v1.3.5 maintenance updates, through 2026-02-02)¶
At a glance
/v1/images/edits— OpenAI-compatible image editing backed by Amazon Bedrock./v1/images/variations— OpenAI-compatible image variations.DEFAULT_TTS_LANGUAGE(v1.3.2) — a configurable default language for text-to-speech, plusimage[]array-style notation for image edits.- Tool-call and streaming fixes (v1.3.1, v1.3.3–v1.3.5) — robust JSON parsing and validation of tool arguments, no premature
contentBlockStopin streamed chat completions, and empty content blocks skipped in assistant responses. - A deprecation mapping (v1.3.4) —
amazon.titan-image-generator-v2:0toamazon.nova-canvas-v1:0.
Adds support for OpenAI's image editing and variation endpoints, enabling image manipulation capabilities backed by Amazon Bedrock. Includes maintenance updates for content block handling, tool call validation, streaming fixes, and TTS optimization.
Image Generation¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/images/edits | ||
/v1/images/variations |
Speech & Audio (v1.3.2)¶
| Feature | Description |
|---|---|
DEFAULT_TTS_LANGUAGE setting | Configurable default language for TTS to optimize performance |
Fixes & Maintenance (v1.3.1–v1.3.5)¶
v1.3.5
- Refactor content block handling to skip empty entries in assistant responses
v1.3.4
- Handle invalid tool call arguments with robust JSON content validation
- Add deprecation mapping for
amazon.titan-image-generator-v2:0→amazon.nova-canvas-v1:0
v1.3.3
- Remove premature stop condition for
contentBlockStopin streaming chat completions
v1.3.2
- Support
image[]array-style notation for OpenAI image edits - Handle empty audio segments in transcription duration calculation
v1.3.1
- Improve JSON parsing for tool arguments and results
- Correct
example→examplesin OpenAPI model path parameter
v1.2.0 – 2025-12-18 – Service Tiers, System Tools & Performance Enhancements¶
At a glance
- Service tiers — the
service_tierparameter on chat completions, with latency headers on all Bedrock routes. - Bedrock-specific system tools — Amazon Nova grounding on
/v1/chat/completions. - GPT5.2 API update —
reasoning_effort=xhighaccepted. - A configuration flag for guardrail override allow, controlling what a request may override on Amazon Bedrock Guardrails.
- Python 3.14 with performance optimization, and direct
aiobotocoreusage replacingaioboto3.
Introduces service tiers and latency headers for all Bedrock routes, Bedrock-specific system tools (Nova grounding), GPT5.2 API compatibility, configurable guardrail overrides, and Python 3.14 optimization.
Chat Completions¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/chat/completions service_tier parameter | ||
/v1/chat/completions Bedrock-specific system tools (Nova grounding) | ||
/v1/chat/completions GPT5.2 API update (reasoning_effort=xhigh) |
Content Safety & Moderation¶
| Feature | AWS Backend |
|---|---|
| Configuration flag for guardrail override allow |
Platform Features¶
| Feature | AWS Backend / Description |
|---|---|
| Service tiers and latency headers (all Bedrock routes) | |
| Python 3.14 support | Upgraded to Python 3.14 with performance optimization |
| Dependency update | Direct aiobotocore usage (replaced aioboto3) |
Fixes¶
- Fix warnings for duplicated FastAPI routes (
/docsand/openapi.json).
v1.1.0 – 2025-11-27 – Embeddings Enhancement, Prompt Caching & Advanced Routing¶
At a glance
- Multimodal embeddings — Amazon Nova multimodal models and TwelveLabs Marengo V3.
- Intelligent embedding plumbing — S3 multimodal upload and sync/async Bedrock invocation chosen per request.
- Prompt caching —
prompt_cache_keyon/v1/chat/completions, plus the GPT5.1 API update (reasoning_effort=none). - Advanced routing — Amazon Bedrock application inference profiles and prompt routers.
- ARN handling — server-side ARN mapping, with optional client-side ARN passing.
Expands multimodal embedding capabilities, adds prompt caching support, and introduces advanced routing with application inference profiles and prompt routers.
Chat Completions¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
Prompt caching /v1/chat/completions prompt_cache_key | ||
/v1/chat/completions GPT5.1 API update (reasoning_effort=none) |
Embeddings¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
| Intelligent S3 multimodal upload | ||
| Intelligent Sync/async Bedrock invocation | ||
| Multimodal embeddings models | ||
| Marengo V3 models |
Advanced Routing¶
| Feature | AWS Backend |
|---|---|
| Application inference profiles | |
| Prompt routers | |
| Server-side ARN mapping | |
| Client-side ARN passing (optional) |
Fixes¶
/v1/chat/completions: Fix default value passed to the converse API for tools without parameters.- stdapi-ai Terraform module: Fix error if alarms_enabled = true but sns_topic_arn undefined.
v1.0.0 – 2025-11-10 – Foundation Release¶
At a glance
- Chat completions —
/v1/chat/completionson every model supporting the Converse and ConverseStream APIs, with DeepSeekreasoning_contentand Qwen thinking parameters. - Embeddings —
/v1/embeddingson Cohere Embed V3 and V4, TwelveLabs Marengo V2, and Amazon Titan Embed V1 and V2. - Speech and audio —
/v1/audio/speech,/v1/audio/transcriptionsand/v1/audio/translations, on Amazon Polly, Amazon Transcribe and Amazon Translate. - Image generation —
/v1/images/generationson Amazon Nova Canvas, Amazon Titan Image Generator and Stability AI models. - Platform — Amazon Bedrock Guardrails, cross-region inference and multi-region failover, Amazon S3 file storage, static token authentication from SSM Parameter Store or Secrets Manager, AWS X-Ray tracing and CloudWatch structured logging.
The initial release establishes core OpenAI API compatibility with Amazon Bedrock backing.
Chat Completions¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/chat/completions | ||
| All models supporting Converse/ConverseStream APIs | ||
/v1/chat/completions reasoning_content | ||
enable_thinking + thinking_budget parameter | ||
top_k parameter |
Embeddings¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/embeddings | ||
| Embed V3 & V4 models | ||
| Marengo V2 models | ||
| Embed V1 & V2 models |
Speech & Audio¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/audio/speech | ||
/v1/audio/transcriptions | ||
/v1/audio/translations |
Image Generation¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/images/generations | ||
| Canvas V1 models | ||
| Image Generator V1 & V2 models | ||
| Image Core, Ultra and SD3.5 Large models |
Model Discovery¶
| Provider | Endpoint/Feature | AWS Backend |
|---|---|---|
/v1/models |
Platform Features¶
| Feature | AWS Backend |
|---|---|
| Bedrock Features | |
| Content filtering and safety | |
| Cross-region inference | |
| Application inference profiles | |
| Model parameters (temperature, top_p, etc.) | |
| Multi-region failover | |
| Bedrock guardrails | |
| AWS Services | |
| File storage | |
| Authentication | |
| Static token authentication | |
| Development mode (no auth) | |
| Observability | |
| Distributed tracing | |
| Structured logging | |
| Health check endpoint | |
| HTTP/Security | |
| CORS support | |
| Trusted host validation | |
| Proxy headers (X-Forwarded-*) | |
| GZip compression | |
| 📚 Documentation | |
| Interactive API docs & OpenAPI schema | |
| 🔌 Compatibility | |
| Provider-specific parameters |