Skip to content

Releases & Roadmap

stdapi.ai is under active development with regular feature releases.

Version Index

Latest: v1.19.1 — released 2026-09-29.

Every release, newest first. Each entry in the Release History below opens with a five-bullet summary.

Version Date Theme Release notes
v1.19.0 (and v1.19.1) 2026-09-23 (2026-09-29) Long conversations, token counting on every model, reasoning control, medical transcription Read
v1.18.0 2026-09-19 Vendor parity across every mirrored API, tenant key rotation, per-tenant rate limits, realtime tools Read
v1.17.0 2026-09-08 Your own model endpoints, the Ollama dialect, per-tenant keys, usage and cost from the API, WebRTC Read
v1.16.0 (and v1.16.1) 2026-08-21 (2026-08-25) Conversations, batches, vector stores, realtime speech and per-user identity Read
v1.15.0 2026-08-03 Reliability, performance and feature completeness Read
v1.14.0 2026-07-12 Bedrock Mantle, video generation, Cohere APIs, moderation and stored conversations Read
v1.13.0 2026-07-03 Terraform module compliance and security hardening Read
v1.12.0 2026-05-29 Completions API, video understanding and file references Read
v1.11.0 (through v1.11.4) 2026-05-02 (2026-05-28) MCP server, agent discovery and model search Read
v1.10.0 2026-04-17 OpenAI Responses API Read
v1.9.0 2026-04-10 Files API and a JSON body for the Images API Read
v1.8.0 2026-04-04 Broader model compatibility and structured output Read
v1.7.0 2026-03-20 Automatic region routing, deprecated model fallback and resilience Read
v1.6.0 2026-02-27 Anthropic API compatibility and advanced Claude capabilities Read
v1.5.0 (and v1.5.1–v1.5.2) 2026-02-15 (2026-02-18) Advanced reasoning and model compatibility Read
v1.4.0 2026-02-11 Audio enhancements and model compatibility Read
v1.3.0 (through v1.3.5) 2026-01-11 (2026-02-02) Image editing and variation support Read
v1.2.0 2025-12-18 Service tiers, system tools and performance Read
v1.1.0 2025-11-27 Embeddings, prompt caching and advanced routing Read
v1.0.0 2025-11-10 Foundation release Read

Roadmap (Tracked on GitHub)

Pending features and current deployment state are tracked on the GitHub Project.


Release History

v1.19.0 – 2026-09-23 – Long Conversations, Reasoning Control & Medical Transcription (with v1.19.1 maintenance update, 2026-09-29)

At a glance

  • Long conversations keep going. The Responses API serves truncation: "auto", context_management compaction and the compaction_trigger item, where all three answered 400 or were dropped. On Anthropic Messages, Claude clears old tool results and thinking blocks through context_management and reports what it cleared.
  • A full context window is refused the way the vendor refuses it. 400 context_length_exceeded on the OpenAI routes, prompt is too long on Anthropic Messages, before the first byte of a stream, and never in the backend's own words. A malformed request is refused the way the same vendor endpoint refuses it, with the parameter and code where the vendor names them.
  • Token counting answers for every text model. input_tokens and count_tokens count past the context window, with server tools, images and documents, on every model including those served by Bedrock Mantle and your own endpoints. Claude up to 4.6 is counted exactly; the others get an approximation that errs high, never low.
  • Reasoning does what the request asks. The effort level now reaches OpenAI GPT-5.x, GPT-6, gpt-oss and Moonshot Kimi K3, which were dropping it; recent Claude models return their reasoning when asked for a summary; and Claude Code's interleaved thinking is no longer filtered out.
  • Medical transcription. amazon.transcribe-medical transcribes clinical dictation and patient–clinician conversations in US English, streamed phrase by phrase, with six medical specialties.

This release is about what happens when a conversation gets long. An agent session eventually outgrows its model's context window, and the gateway had one answer for that: a refusal, often in the backend's words rather than the API's. The Responses API now trims or compacts a conversation when asked, Claude clears stale tool results through Anthropic's context editing, an overflow that still happens is refused exactly as the vendor refuses it, and a client can count tokens on every text model to see it coming. Around that: reasoning controls that now reach the model families that were dropping them, GPT-6 web search and code interpreter with no configuration, Moonshot Kimi K3, medical transcription, and a pass over reported costs that corrects several prices, one of them reported up to 50% above what AWS bills.

New IAM Permissions, All Optional

Nothing new is required. A deployment that does not serve medical transcription needs the policy it already has; the Terraform module grants these wherever it grants transcription. See Speech-to-Text.

  • Medical transcription: transcribe:StartMedicalTranscriptionJob and transcribe:StartMedicalStreamTranscription on * (neither accepts a resource type), transcribe:GetMedicalTranscriptionJob and transcribe:DeleteMedicalTranscriptionJob on medical-transcription-job/*, and transcribe:TagResource extended to that same resource. Without them amazon.transcribe-medical is unavailable, and standard transcription is unaffected.

Behavior Changes

Review these before upgrading. They may change what existing clients observe:

  • The OpenAI GPT-6 family is served by Bedrock Mantle by default, wherever a configured Mantle region lists it, so its web_search and code_interpreter work with no configuration. AWS_BEDROCK_MANTLE_PREFERRED_MODELS now defaults to openai.gpt-5.6,openai.gpt-6. It is the trade GPT-5.6 already made: no Global routing discount, so GPT-6 Astra costs exactly 10% more per token, no guardrails, and usage billed and reported under Bedrock Mantle. An empty value keeps both families on the classic endpoint. An explicit value replaces the default, so a deployment that already sets one keeps GPT-6 where it was.
  • A context-window overflow is refused in the API's own terms. Chat Completions answers 400 context_length_exceeded naming messages, Responses the same naming input, Anthropic Messages invalid_request_error prompt is too long: N tokens > M maximum, and legacy Completions and Ollama their own wording, on every model, Bedrock Mantle included. A streamed Chat Completions, Completions, Messages or Ollama request that the model refuses on its first event now gets that 400 before any byte, instead of a 200 followed by an error event. A streamed Responses request keeps the in-stream error then response.failed, as OpenAI streams it, and that error event now also nests the error under error, as the live API does, so the official SDK raises it.
  • Validation errors are worded as each vendor words them. On the OpenAI routes error.param names the field in OpenAI's notation (messages[0].content) and error.code the failure (unknown_parameter, missing_required_parameter, invalid_type, invalid_value and the range codes); Anthropic routes answer <field>: <reason> with no gateway prefix. A client matching on the old message text needs updating. See Using stdapi.ai.
  • STRICT_INPUT_VALIDATION refuses an unknown field only where OpenAI does. It is now ignored on chat messages other than developer, on content parts, tool calls and tool definitions, on Responses function tools and on moderation requests. A replayed tool call carrying an unknown field is now ignored whatever the setting.
  • Token counting no longer answers 400 for any text model. Models served by Bedrock Mantle (the GPT-5.6 and GPT-6 families by default), Marketplace and SageMaker AI endpoints, inputs past the context window and requests offering server tools are all counted, and search_models lists both counting routes for every text model. Most of these counts are approximate: see Known Limitations below.
  • reasoning.summary on the Responses API now returns a summary. Any value puts the reasoning in the item's summary as summary_text parts, streamed as response.reasoning_summary_* events, where it was ignored and the text left in content. Codex, which displays summary alone by default, now shows the reasoning.
  • A requested reasoning effort now takes effect on OpenAI GPT-5.x, GPT-6 and gpt-oss, where it was silently dropped, so their output-token usage follows the level you ask for rather than the model's default. minimal, which none of them accepts, is sent as low.
  • Transcription usage is recorded to the second, with no 15-second minimum, matching AWS's published billing: a 2-second clip was recorded as 15. A streamed transcription is costed at the streaming rate and carries input_seconds_by_spec: {"streaming": n}, where the batch rate under-reported it by 40% in us-east-1.
  • A streamed transcription no live session can serve falls back to a transcription job rather than answering 503. When no configured region opens a live session, for lack of the transcribe:StartStreamTranscription permission or of a region that streams the model, the request is still streamed, but its events arrive together at the end and it needs a transcription bucket.
  • A bare computer tool on Claude Opus 5 and 5.5 is now the computer-use server tool, as on the other current Claude models, where it was passed through as a custom tool.
  • Kimi K3 is no longer listed as Batch API capable: AWS refuses batch inference for it.
  • Compaction item IDs now start with cmp_, as upstream's do.

Known Limitations

  • Most token counts are approximate, and err high. A count is exact for Claude up to 4.6 and for the Claude models served by Bedrock Mantle. Claude 4.7 and later (Opus 4.7, Opus 5.5, Sonnet 5, Fable) are counted 7% to 70% above the exact figure, about 50% typically. Every other model, the Mantle-served GPT families and your own endpoints included, is counted at least 5% above it and about 30% typically, more for one-line prompts and non-Latin scripts. Across 27 tokenizers and everything measured, no approximation came in below the exact count. On those models compact_threshold is approximate too, and truncation: "auto" counts the whole input. A file given by URL, S3 URI or file ID is counted from its size and never downloaded. The usage a response reports is always exact. See Input Token Counting.
  • An exact count can drift slightly past the context window and on a history holding web search results: within 20 tokens on a 300,000-token message, and within 2% with search results, on Claude Haiku 4.5.
  • Recent Claude models return no reasoning text on Chat Completions. Opus 4.7 and later, Sonnet 5, Fable and Mythos return it only as a summary, which Chat Completions cannot request: use thinking.display on Messages or reasoning.summary on Responses.
  • Context editing is Claude-only, and has no compaction. Other models accept context_management and ignore it with a server-log warning, compact_20260112 is refused as an unknown edit type, and batch results do not report applied_edits.
  • Compaction items are not encrypted. They carry the conversation's text, system and developer messages included: hand them only to parties that may read and alter that conversation, or continue with previous_response_id to keep it on the server.
  • GPT-6 Sol and Luna have no published AWS price yet, so their usage is recorded without a cost. Mantle lists GPT-6 in few regions: where no configured Mantle region does, web_search and code_interpreter are refused with 400.
  • Medical transcription is US English only, refuses srt and vtt (verbose_json carries timed segments), and takes a specialty other than primary care only when streamed on a deployment serving live medical transcription.

New API Features

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI truncation: "auto" on Responses – an input over the context window has its oldest turns dropped, then the latest turn's longest text cut at its end, until it fits. System and developer messages are always kept, usage reports the input that was kept, and input_tokens counts what a response would keep Amazon Bedrock Amazon Bedrock - Converse API, Bedrock Mantle
OpenAI OpenAI context_management compaction on Responses – once the input crosses compact_threshold, everything but the latest step is summarized into a compaction item that leads output and stands for the whole history when sent back; streamed at index 0 before the answer. A trailing compaction_trigger item compacts on demand, where it was silently dropped. Both calls are billed, and usage adds them up Amazon Bedrock Amazon Bedrock - Converse API, Bedrock Mantle
OpenAI OpenAI & Anthropic Anthropic The output is capped when only it overflows – an input the window holds alone but not beside max_output_tokens or max_tokens is answered with a shorter output on Responses and Anthropic Messages, as both vendors answer it. Chat Completions keeps refusing it, as OpenAI does Amazon Bedrock Amazon Bedrock
OpenAI OpenAI & Anthropic Anthropic Token counting on every text model – /v1/responses/input_tokens and /v1/messages/count_tokens count an input past the context window, always above the window so a client learns the conversation no longer fits, plus server tools with their replayed calls, results and citations, inline images, and documents and remote files at an allowance Amazon Bedrock Amazon Bedrock
Anthropic Anthropic Context editing – clear_tool_uses_20250919 and clear_thinking_20251015 with every documented option. Claude applies the edits and reports them in applied_edits, on the final message_delta when streaming, and count_tokens counts the edited prompt beside original_input_tokens. The beta flag is added for you, so the header is optional Amazon Bedrock Amazon Bedrock - Converse API, Bedrock Mantle
Anthropic Anthropic thinking.display – summarized or omitted reaches Claude as sent, so Opus 4.7 and later, Sonnet 5, Fable and Mythos, which omit their thinking text by default, return a summary when asked. The thinking tokens bill the same either way. On Responses, reasoning.summary asks for the same summary Amazon Bedrock Amazon Bedrock - Converse API
OpenAI OpenAI Medical transcription – amazon.transcribe-medical on /v1/audio/transcriptions, for a physician's dictation or a patient–clinician conversation (Type), with Specialty and medical custom vocabularies. json, text, verbose_json and diarized_json are served, and stream=true streams phrase by phrase. A client that can only send a model name reaches a fixed variant through a model alias Amazon Transcribe Amazon Transcribe Medical

Models & Reasoning

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI GPT-6 web search and code interpreter – served with no header and no configuration wherever a configured Mantle region lists the model, and Chat Completions reaches GPT-6 natively there rather than converted Amazon Bedrock Amazon Bedrock Mantle
OpenAI OpenAI Reasoning effort on GPT-5.x, GPT-6 and gpt-oss – from every dialect: reasoning_effort, reasoning.effort, output_config.effort or thinking. none turns reasoning off where the model can; GPT-6 Astra, and gpt-oss outside Bedrock Mantle, cannot, and are served at their default level with a warning in the request log. gpt-oss gets its own low / medium / high scale Amazon Bedrock Amazon Bedrock
Moonshot Moonshot Kimi K3 – reasoning effort from every dialect, with none answering without reasoning, and automatic prompt caching: a repeated prefix is reported as cached tokens on Chat Completions, Responses and Messages and costed at the cache rates. Cache breakpoints are accepted and ignored, since the model rejects them. A turn replaying the model's own reasoning is served Amazon Bedrock Amazon Bedrock
Anthropic Anthropic Claude Opus 5.5 – a request disabling reasoning on Opus 5.5, which always reasons, is served at its adaptive default with a warning in the request log, as on Fable and Mythos, rather than refused. On Bedrock Mantle, xhigh and max reach Claude 4.7 and later unchanged instead of being capped to high Amazon Bedrock Amazon Bedrock

Cost Reporting

  • Daybreak Blue (GPT-5.6 Sol) was reported above what AWS bills: $5.50 / $33.00 per million input / output tokens against AWS's $4.40 / $22.00, and $11.00 / $49.50 against $8.80 / $33.00 past 272K, with the cache rates 25% high as well. It is now costed at GPT-5.6 Sol's rates.
  • GPT-5.4 and GPT-5.5 are costed at their long-context rates past 272K input tokens, the boundary AWS publishes for them, where every call was costed at the short rate. See Long-Context Pricing.
  • GPT-6 Astra is priced: In-Region $11.00 / $55.00 and Global $10.00 / $50.00 per million input / output tokens, with cache and long-context rates. It had no price at all.
  • Kimi K3 calls are priced, where every call reported no cost. A Global call is costed at its Global rate ($3.00 / $15.00) rather than the US one ($3.30 / $16.50), and cache reads and writes are priced. A Grok 4.6 Global call now reaches its Global rate ($2.00 / $6.00) the same way.

Platform Features

  • Transcription moves on from a region that does not offer the operation. Medical transcription and live streaming run in fewer regions than standard transcription; a region lacking one now hands the request to the next configured region instead of failing it, and when none offers it the answer is a 503. See Resilience.
  • The Models page lists Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna (the last two without a price, as AWS publishes none yet), GPT-6 Astra with its prices, the GPT-5.4 and GPT-5.5 long-context tier, and the Global rates of Kimi K3 and Grok 4.6.

Fixes

  • Claude Code's interleaved thinking reaches Claude. The allowlist carried the interleaved-thinking flag capitalised, which no client sends, so thinking between tool calls silently stayed off. Separately, a request carrying the memory tool or a computer-use tool lost every flag of its anthropic-beta header, the 1M context window's included, when the gateway added the one that tool needs; the two lists are now joined, and count_tokens merges and filters the flags the same way.
  • Reasoning effort on Claude 3.7 to 4.5 served by Bedrock Mantle no longer fails: it was sent as an effort level these models refuse, so every such request answered 400. It now becomes the thinking budget they take, scaled on the output limit; an output limit of 1,024 tokens or less leaves no room for one, and the request is served without reasoning and a logged warning.
  • A context-window overflow no longer leaks the backend's error text, which on some models named an internal service and a request identifier, and a streamed one is no longer reported as a retryable server_error.
  • A failed Responses request on a Bedrock Mantle model converted to that API is refused, where it was answered as an empty success.

Fixes & Maintenance (v1.19.1)

  • Claude Sonnet 5.5 is served without errors. Turning reasoning off or forcing a tool surfaced a raw 400 from Amazon Bedrock. Thinking turned off on any route (reasoning_effort: "none", enable_thinking: false, reasoning.effort: "none", thinking: {"type": "disabled"}, think: false) now reaches it as between_tools, its lowest setting, and a forced tool choice is refused with a message naming auto, as on Opus 5.5. See Turning Thinking Off on Claude Sonnet 5.5.
  • Anthropic Messages accepts thinking: {"type": "between_tools"}, as Anthropic does, on message creation and token counting.
  • thinking: {"type": "disabled"} beside an output_config.effort turns thinking off and keeps the effort, as upstream does. Claude reasoned at that effort instead, and was billed for it.
  • Grok 4.7 calls are priced, at Global $2.00 / $6.00 per million input / output tokens. Every call was reported as free.
  • GPT-6 Astra short prompts are costed at the short-context rate. The Price List rows AWS now publishes for it were read as one tier, so a short prompt could be costed at the long-context rate, up to twice as high.
  • GPT-6 Luna, GPT-6 Sol and GLM 4.6 are priced, at the rates on their AWS model cards (GPT-6) and on Z.ai's pricing page (GLM 4.6). Every call was reported as free.
  • The Terraform module grants the server the KMS permission its DynamoDB table needs. Since module 1.17.0, a deployment where the module created the table refused every table read and write, so tenant API keys, per-tenant rate limits and the shared model cache were unavailable. Update the module to apply it.
  • A conversation replaying a turn that ended after reasoning is answered, where automatic prompt caching made Amazon Bedrock refuse it with a 400; Codex on Nova 2 Lite could not recover from a cut-short answer.
  • The Models page lists Claude Sonnet 5.5 and Grok 4.7, names each model once where it is served by two AWS services, quotes the nearest region's rate where the selected region has none, lists Amazon Transcribe in every region that offers it, shows closed models' licence as proprietary, and takes each model's APIs and prompt caching from its AWS model card.

v1.18.0 – 2026-09-19 – Vendor Parity, Tenant Key Rotation & Rate Limits

At a glance

  • Requests the real APIs accept are no longer refused — about thirty of them, across Chat Completions, Responses, Anthropic Messages, the Ollama dialect, audio and images: a temperature above 1, service_tier: "fast", max_tokens: 0, speed up to 4.0, an empty Ollama format, a tool_result with no content, and the rest. The gateway's job is to answer what the vendor answers.
  • Tenant keys rotate themselves — set TENANT_KEY_SECRETSMANAGER_PREFIX and each key becomes a Secrets Manager secret the tenant reads for itself, rotated on a schedule with an overlap window so nothing is locked out mid-flight.
  • Per-tenant rate limits — requests_per_minute and tokens_per_minute per key or per deployment, counted across every instance. A deployment that declares none pays neither the latency nor the DynamoDB bill.
  • Voice agents can act — a Realtime session declares tools and answers the calls the model makes, and Chat Completions requests can ask for a server-side web search with OpenAI's own web_search_options.
  • Uploads stream — an upload part is handed straight to S3 instead of being held in memory, so the per-part ceiling rises from 64 MiB to S3's own 5 GiB and a whole upload from 8 GiB to 48.8 TiB.

This release is about being the API it claims to mirror. The bulk of it is a compatibility sweep run against the vendors' own endpoints and official clients, not just against the gateway: every divergence it found is either fixed here or documented as a limitation. Alongside it, three things an operator asked for — tenant keys that rotate without anyone moving a secret by hand, per-tenant request and token limits that hold across instances, and a guardrail cost lever that scopes evaluation to the turns you name. Two capabilities round it out: tools on the Realtime API, which is what a voice agent needs to do anything beyond talking, and web search on Chat Completions, the dialect most clients actually speak.

New IAM Permissions, All Optional

Nothing new is required. A deployment that enables neither feature below needs the policy it already has; the Terraform module grants each set when its feature is turned on. See IAM Permissions.

  • Tenant key rotation — five Secrets Manager actions scoped to <prefix>/*, plus kms:GenerateDataKey and kms:Decrypt conditioned on the service and the encryption context, for TENANT_KEY_SECRETSMANAGER_PREFIX. They replace the ssm:* delivery actions rather than adding to them. See Tenant Key Rotation.
  • Tenant rate limits — dynamodb:UpdateItem on the shared table, in its own statement and conditioned on the counter items, so the one write-in-place action cannot touch a tenant record. Granted only when a limit is declared. See Shared Table.

Behavior Changes

Review these before upgrading — they may change what existing clients observe:

  • An upload part is no longer capped at 64 MiB, and an upload is no longer capped at 8 GiB. Parts are streamed to S3 rather than buffered, so S3's limits are the only ones left: 5 GiB per part, 10,000 parts, 48.8 TiB per object. This leaves the gateway more permissive than OpenAI, which caps every part at 64 MiB — a client written to the vendor's limit is unaffected. A part sent inline as JSON keeps the 64 MiB bound, because that form is decoded whole and the number is a memory bound there. Large parts cost ephemeral disk rather than memory: the request body is spooled before it is streamed out.
  • A phase on a non-assistant Responses input message is now refused. The official client types phase under every role; the API keeps it under assistant alone and answers 400 unknown_parameter for the others. The gateway accepted it silently and now answers what the vendor answers. A null phase is not a refusal.
  • tool_choice: "none" is now documented as Partial, not Supported, on all three compatibility tables. The tools a turn declares are the only ones the model may call — that part was already true and is now proven by tests. What the backend cannot express is a conversation that has already called a tool: declaring no tool, or tool_choice: "none", leaves that conversation's tools callable. Nothing changed in the code; the tables stopped claiming otherwise.
  • An Anthropic file listing now always carries next_page, null on the last page, and the envelope no longer drops its null fields. This is what anthropic.pagination.SyncPageCursor reads; before it, a client paging with the SDK saw one page and stopped. The after_id and before_id cursors still work and still page in both directions.
  • Three Bedrock responses now report what actually happened rather than what was asked for: the OpenAI surfaces report the service tier that served the request, a streamed Chat Completions response carries the same moderation results the non-streamed one returns, and a Responses object carries the truncation the request asked for.
  • The validation-error log line no longer contains the request body. A 422 is logged as the field paths that failed validation, without their values, because a body can carry a credential. The line stays debuggable; there is deliberately no setting to restore full bodies.
  • A queued vector-store indexing job is refused for a tenant-credential key, rather than running that tenant's embeddings on the deployment's AWS account. This matches the refusal batch jobs already carry, and is a documented limitation.
  • A knowledge-base vector store served search-only refuses corpus deletion. Any caller could previously delete documents from the underlying knowledge base through a store that was meant to read it.
  • An expired Files API object is refused when referenced by file_id. Expiry was honoured on every other path; a file_id input bypassed it and kept feeding inference until the storage lifecycle rule swept the bytes.
  • A file of a type a vector store cannot index is refused by the attach, with a 400 naming what that store does index, where it was previously accepted and settled failed with unsupported_file a moment later. Image, audio and video types are now unindexable on their content type alone, as the archive and office types already were, which is what the official API answers. A file attached in a batch is still reported on the file, as a batch reports everything that fails in it, and a file whose bytes turn out not to be what its content type claims still settles failed.

New API Features

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI Realtime function calling – a session declares tools and answers the calls the model makes, so a voice agent can do more than talk Amazon Bedrock Amazon Bedrock - Amazon Nova Sonic
OpenAI OpenAI web_search_options on Chat Completions – OpenAI's only spelling for server-side web search on the dialect most clients speak, resolved through the same declaration the Responses and Messages dialects already use. A model that runs no web search refuses rather than answering from its own knowledge, which nothing in the response would distinguish Amazon Bedrock Amazon Bedrock - Nova Grounding, Anthropic web search
OpenAI OpenAI Prompt-attack and personal-data moderation – jailbreak and prompt-injection detection, and personal data detection, on /v1/moderations and the moderation parameter of the chat routes, with no guardrail resource to configure Amazon Bedrock Amazon Bedrock Guardrails
OpenAI OpenAI echo on /v1/completions – the prompt is echoed back ahead of the completion, which the gateway can do locally and was silently dropping Amazon Bedrock Amazon Bedrock - Converse API
Anthropic Anthropic Opaque page tokens on GET /v1/files – next_page is served and page accepted, the way the SDK now pages. The token is opaque and refused when it is not one this API issued; the ID cursors stay Amazon S3 Amazon S3
Anthropic Anthropic expires_in_seconds on a file upload – the full upstream range (1 hour to 90 days) is honoured and expires_at reported on all three routes. Expiry is enforced on every read, so a TTL longer than the storage sweep still expires at the instant promised Amazon S3 Amazon S3
Ollama Ollama prompt_eval_cached_count, and the base URL's connection test – the cached-prompt count Bedrock already reports is passed through, and the base URL answers Ollama is running rather than a 404. /api/blobs/{digest} joins the store verbs this deployment refuses Amazon Bedrock Amazon Bedrock - Converse API

Multi-Tenancy

Provider Endpoint/Feature AWS Backend
stdapi.ai Tenant key rotation – with TENANT_KEY_SECRETSMANAGER_PREFIX each key is stored as a Secrets Manager secret the tenant reads for itself, so TENANT_KEY_ROTATION_DAYS rotates it with nobody in the loop. The superseded key keeps working for TENANT_KEY_ROTATION_OVERLAP_SECONDS — seven days by default, 0 for a hard cutover — and only the immediately previous key survives AWS Secrets Manager AWS Secrets Manager
stdapi.ai Per-tenant rate limits – requests_per_minute and tokens_per_minute, per key or as a deployment default, counted in the shared table so a limit holds across every instance. Tokens are admitted on an estimate and reconciled from the billed usage, which is what keeps a Realtime or MCP session from billing to a window that has ended. A counter that cannot be written fails closed with a 503, never a 429 the client did not earn Amazon DynamoDB Amazon DynamoDB

Cost Control

Provider Endpoint/Feature AWS Backend
stdapi.ai AWS_BEDROCK_GUARDRAIL_SCOPE_TURNS – a guarded conversation is charged for everything the client replays, because Converse evaluates the whole history on every turn. This scopes input evaluation to the trailing user turns you name: a 6,401-character history bills 7 text units per policy and adds 409 ms, against 1 unit and 194 ms for the last turn alone. Off by default and left off unless you decide otherwise — scoping genuinely lowers detection, not just the bill: a prompt attack split across two turns is blocked today and is not blocked once only the last turn is tagged. Model output is always evaluated in full Amazon Bedrock Amazon Bedrock Guardrails

Platform Features

  • Uploads stream to S3. A binary part is handed from the request straight to S3 and re-read in 8 MiB chunks for the session digest, so memory is one chunk whatever the part size. The digest keeps every guarantee it had: the fold still runs under the session lock inside a shield, and the signature is still written in the same critical section as the bytes, so a digest that does not answer for exactly the parts its signature names still cannot reach a completion.
  • A batch input file is streamed end to end, rather than being materialised several times over — decoded lines, parsed bodies and a joined upload blob could each hold up to 200 MiB of one request.
  • The documentation pages ship a current Swagger UI and ReDoc — swagger-ui-dist 5.33.0, which lays out only the part of a large specification that is on screen, and redoc 2.5.4, which carries accessibility fixes and patched bundled dependencies. /redoc no longer loads anything from cdn.redoc.ly, which the documentation had always promised it did not.

Fixes

  • A tiny multipart field name can no longer exhaust the gateway: a bracketed index in a form field name drove an unbounded allocation, so a field a few bytes long could OOM-kill the instance. The index and the nesting depth are now bounded, above what every official client actually sends.
  • An application-inference-profile ARN no longer hijacks a model's routing: resolving a model from an ARN copied the catalogue entry shallowly and then wrote through the copy, so one caller's ARN changed where every later caller in that region was routed. The lookup now copies what it mutates.
  • A realtime session opened with an ephemeral client secret is guarded: the deployment's guardrail was configured on every other entry point and skipped on that one, so the untrusted-browser credential the guardrail exists for was the one case it did not cover.
  • Batch creation no longer blocks every other caller, and a queued vector-store indexing job no longer strands a file or orphans its vectors when the job is interrupted or the lookup fails.
  • The tenant key reconciliation loop survives an unexpected error: one failure ended it silently, and with it every later mint and rotation, until the instance was restarted.
  • Audio and media stop answering with the wrong thing: a truncated encode was served as a complete 200 when its input failed mid-stream, a bad upload to a live transcription never got its 400, and text with an & or a < produced a Bedrock InvalidSsmlException whenever speed was not 1.0. A transcription job that is already gone is now accepted as deleted rather than re-raising.
  • A file URL that is fetchable is now usable: URL ingestion follows redirects and falls back to a ranged read, where before a redirect or a server refusing a HEAD probe failed the request outright.
  • A stored conversation stays readable: an additional_tools or compaction_trigger item made every later read of that conversation a 500, permanently.
  • Requests the vendors accept are accepted: temperature above 1, service_tier aliases including "fast", max_tokens: 0 (the documented cache pre-warm), the 2026 browser and computer toolsets, the container object form, the custom-voice object, a tool_result block with no content, speed up to 4.0 on the Polly engines that honour it, an explicit JSON null on ten nullable image fields, output_compression: 0, an empty Ollama format (which broke langchain-ollama on every call), and the Ollama option sentinels seed: -1 and top_k: 0 that answered 500.
  • Responses the vendors send are sent: event: names on streamed image frames, the citations a streamed document answer carries, background and safety_identifier on the Converse path, the code_interpreter_call item the model actually ran instead of an erased one, an item reference answered with the item it names in an envelope the official SDK can parse, and the filenames the vendor keeps rather than sanitising.
  • A realtime turn the server detects is now committed like a turn the caller ends: on server_vad, the API's default turn mode, the gateway sent neither input_audio_buffer.committed nor the conversation item the turn became, so a client waiting for either against a detected turn waited for ever.
  • The Anthropic file listing honours ids=, where it previously ignored it and returned the wrong set of files.
  • Ollama chat keeps what the caller sent: images on a tool, assistant or system message were silently dropped, and two endpoints answered an accidental 404.
  • The Models page states what the source measured: two benchmark scores were published against the wrong model — GLM 4.7 carried GLM 4.7 Flash's GPQA Diamond result, and DeepSeek V3.2 carried the previous generation's.
  • Four cells in the features comparison table had drifted from what the compared products now do.
  • Four silent limitations are now written down: a Chat Completions message name is accepted and ignored, the verbose_json transcription decoder fields are placeholders rather than measurements, a bad batch input file is reported per line rather than refusing the batch, and the moderations input cap is the gateway's own rather than a vendor rule the docs implied it was.

v1.17.0 – 2026-09-08 – Your Own Models, Your Own Tenants, Your Own Spend

At a glance

This release widens what a deployment can serve, who it can serve it to, and what it can tell you about the bill. Models you host yourself: Amazon SageMaker AI endpoints and Amazon Bedrock Marketplace model endpoints you have already deployed join the model list and answer on the same routes as everything else — a fine-tuned model, a model no serverless catalogue carries, or one you scale to zero between requests, reached with the same client code. A new dialect: clients written for the Ollama API, and the many tools that speak only that, now reach Amazon Bedrock unchanged. Per-tenant API keys: one deployment serves several tenants, each scoped to the models and endpoints it may use, revocable on its own, and free to bring its own AWS account. Usage and spend, from the API: the OpenAI Administration API reports what was used and what it cost — per model, per endpoint and, where a deployment identifies callers, per user — read from the metrics the gateway itself writes rather than estimated. Speaking to it, live: the Realtime API gained the WebRTC transport, so a browser negotiates its own media session and the audio path terminates in the gateway, with no WebSocket relay in between.

New IAM Permissions, All Optional

Nothing new is required. A deployment that enables none of the features below needs the policy it already has; the Terraform module grants each set when its feature is turned on. See IAM Permissions for the statements in full, and note that a denial now names the action it needs in the server log rather than leaving you to guess it.

  • Your own SageMaker AI endpoints — sagemaker:CallWithBearerToken and sagemaker:InvokeEndpoint on the endpoint ARNs you serve, for Amazon SageMaker AI endpoints.
  • Your own Marketplace model endpoints — bedrock:ListMarketplaceModelEndpoints and bedrock:GetMarketplaceModelEndpoint to discover Marketplace model endpoints, sagemaker:InvokeEndpoint and sagemaker:InvokeEndpointWithResponseStream (called via Bedrock on the server's behalf) to invoke them, and — for a deployment with per-user roles — bedrock:InvokeModel on arn:aws:bedrock:*:ACCOUNT_ID:marketplace/model-endpoint/*, which a foundation-model wildcard does not cover.
  • Shared records — the dynamodb item actions on one table, for per-tenant API keys or the shared model list.
  • Tenants — ssm:PutParameter and ssm:GetParameter to deliver each tenant its key once, kms:Encrypt, kms:Decrypt and kms:GenerateDataKey on the key named by TENANT_KEY_SSM_KMS_KEY_ID when you set one, and sts:AssumeRole on exactly the roles tenants declare when they bring their own AWS account.
  • Usage reporting — cloudwatch:GetMetricData and cloudwatch:ListMetrics for the Administration API.
  • Per-user cost attribution roles — one addition to an existing feature: the role the gateway assumes per end user now needs s3:GetObject (and kms:Decrypt through S3) on every bucket a request can reference by URI, because Amazon Bedrock reads a video or document at an s3Location with the invoking identity. Without it, a request that carries stored media answers 403 under that role while the same request under the task role succeeds. The module grants it; a hand-written role needs the statements in Per-User Cost Attribution.

Behavior Changes

Review these before upgrading — they may change what existing clients observe, and what AWS charges you:

  • Three responses now name the model that actually served the request, rather than echoing the string the caller sent. The Anthropic Messages response (/anthropic/v1/messages, streaming and not), the batch object returned by GET /v1/batches/{id}, and the model field inside Anthropic batch result lines — so a request naming claude-opus-5 now comes back naming anthropic.claude-opus-5. This matches what the OpenAI chat, Responses, embeddings and video responses already did, and affects existing callers who use aliases, not only callers using the new wildcard patterns.
  • A price change for the OpenAI GPT-5.6 family. GPT-5.6 Sol, Terra and Luna are now served by Amazon Bedrock Mantle by default, so their web_search and code_interpreter tools work with no configuration — Bedrock serves those on Mantle alone. Mantle has no cross-region inference profiles, so these models stop riding the Global profile and its discount and pay the In-Region rate: exactly 10% more per token, on input, output, cached and long-context rates alike. Per million input / output tokens that is $4.40 / $22.00 for Sol, $2.20 / $13.20 for Terra and $0.22 / $1.32 for Luna, against $4.00 / $20.00, $2.00 / $12.00 and $0.20 / $1.20 before. A deployment that had already disabled AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL pays what it paid. Setting AWS_BEDROCK_MANTLE_PREFERRED_MODELS to an empty value restores the previous routing, and the previous price, for every dual-homed model.
  • A guardrail is never silently skipped on a Mantle-served request. Amazon Bedrock Guardrails do not apply to Amazon Bedrock Mantle, so a model served there under a configured guardrail would answer a caller who asked to be guarded with an unguarded answer, and nothing would report it. Both halves of that are now refused. A deployment that routes a dual-homed model to Mantle under a guardrail — including through a MODEL_ALIASES guardrail — stops at startup, naming the routed models and the way out; because AWS_BEDROCK_MANTLE_PREFERRED_MODELS now has a default, a guardrailed deployment meets this without having changed anything, and clearing that setting restores both the guardrail and the previous routing. A model that only Mantle serves has no such way out: with a guardrail configured — deployment-wide, through MODEL_ALIASES, or named on the request itself — every request to it answers 400, whatever that setting holds.
  • The same rule now covers your own SageMaker AI endpoints. A request to a SageMaker-served model under a configured guardrail answers 400 instead of being served unfiltered, and a MODEL_ALIASES guardrail naming one stops startup. A SageMaker endpoint with no capacity to wake answers 503 (retry later), not 400.
  • Mantle routing and the per-user role's identity requirement are refused together. A Mantle request is signed with the server's own credentials, so AWS_BEDROCK_USER_ROLE_REQUIRE_IDENTITY can never apply to a model routed there: a deployment setting both with AWS_BEDROCK_MANTLE_PREFERRED_MODELS non-empty stops at startup, and a Mantle request without an end-user identity answers 400.
  • The Terraform module requires Terraform or OpenTofu 1.9. An existing deployment has to be on that version before it can upgrade the module.
  • The Terraform module now scales the service on CPU, where before it never scaled at all. autoscaling_cpu_target_percent defaults to 70 instead of creating no scaling policy, so a fleet that has always held one task per availability zone now grows under load up to autoscaling_max_capacity — five times the minimum unless you set it. That raises what the Marketplace license can bill: a 3-AZ deployment metered at $0.10 per container-hour moves from a flat ~$216/month to a range whose ceiling is five times that. The gateway spends most of its time awaiting AWS, so CPU mainly answers the audio and video transcoding; set autoscaling_alb_target_requests_per_target to scale on request volume instead, or autoscaling_cpu_target_percent = null to hold the fleet at a fixed size. A WebRTC media deployment is unaffected: it pins the capacity to one task, and no policy is created when the minimum and maximum are equal.
  • A load balancer with no certificate now says so. alb_enabled without alb_domain_name in a public zone or an alb_certificate_arn serves the API over plain HTTP, which is a supported way to stand up a trial. The plan and apply now carry a warning naming the two variables that turn on HTTPS, rather than leaving it to be noticed. It is a warning, not an error: nothing about the deployment changes.
  • The Terraform module turns its CloudWatch alarms on for a deployment that named a topic. alarms_enabled defaults to whether sns_topic_arn is set, rather than to false: a deployment that gave a topic and never set the flag had no alarms and no notifications, which is the one configuration that cannot have been intended. It now creates the five it was configured for — the four ECS service alarms of the underlying module, plus one on ERROR and CRITICAL log lines — each with its own CloudWatch charge. Setting alarms_enabled = false keeps the previous behaviour, and setting it true with no topic creates the alarms with nothing to notify.
  • A configuration that would silently drop the compliance VPC endpoints is refused at plan time. compliance_vpc_endpoints_enabled and guardduty_vpc_endpoint_enabled need a private application subnet to place their network interface in. Setting nat_gateways_allowed = false while internet access is still required — which it is by default — makes those subnets public, and the endpoints were then simply not created, with nothing said. That combination now fails the plan. A deployment already running it was not getting the endpoints it had asked for: set nat_gateways_allowed = true, or leave it unset, or turn the endpoints off.
  • Input token counting is unavailable for the GPT-5.6 family. POST /v1/responses/input_tokens answers 400 for the models Mantle serves, which now includes that family. Clearing AWS_BEDROCK_MANTLE_PREFERRED_MODELS brings it back.
  • GPT-5.6 usage is reported and billed under Bedrock Mantle, attributed by project rather than by IAM principal, and runs on Mantle's own throughput quotas. Batch inference, prompt caching and response IDs stored before the upgrade are unaffected.
  • The community container image moves from Alpine to a standard Debian base. glibc is what the WebRTC media stack needs — aiortc publishes no musl wheel for any version — so the community image now ships every capability the commercial one does. It declares python as the image entry point and runs as an unprivileged user with uid/gid 65532 instead of 1000. The API, its endpoints and the audio formats it accepts are unchanged, and deployments through the Terraform module are unaffected — the task definition already pinned 65532. Two things change for anyone running the image by hand: arguments after the image name are passed to the Python interpreter rather than run as a program, and a mounted ~/.aws needs --user "$(id -u):$(id -g)" -e HOME=/home/nonroot — or --userns=keep-id:uid=65532,gid=65532 on rootless Podman — to stay readable. It is also a larger image, built from the distribution's own packages; the hardened, minimal one is the AWS Marketplace edition. See Local Development.

Your Own Model Endpoints

Provider Endpoint/Feature AWS Backend
stdapi.ai Amazon SageMaker AI endpoints – the endpoints you already run join the model list under the names you pick and answer on the same routes as every other model, streaming included: a fine-tuned model, or one no serverless catalogue carries, reached with unchanged client code Amazon SageMaker AI Amazon SageMaker AI
stdapi.ai Endpoints scaled to zero – an endpoint with no running capacity is woken and waited for rather than refused; one with no capacity to wake answers 503 (retry later), not 400 Amazon SageMaker AI Amazon SageMaker AI
stdapi.ai Amazon Bedrock Marketplace model endpoints – the endpoints you have deployed are discovered and served like any other model, invoked through Amazon Bedrock on the server's behalf, and one being updated stays served Amazon Bedrock Amazon Bedrock - Marketplace model endpoints

New APIs

Provider Endpoint/Feature AWS Backend
Ollama Ollama POST /ollama/api/chat – conversational responses with tools, images and thinking, streamed as the newline-delimited JSON Ollama clients expect, carrying the timings they read Amazon Bedrock Amazon Bedrock - Converse API
Ollama Ollama POST /ollama/api/generate – a response for a single prompt, streamed or whole; an empty prompt, here or on /api/chat, is the load/unload no-op Ollama defines rather than a 400 Amazon Bedrock Amazon Bedrock - Converse API
Ollama Ollama POST /ollama/api/embed and POST /ollama/api/embeddings – vector embeddings for one or several inputs, including the legacy single-prompt route older clients still call Amazon Bedrock Amazon Bedrock - embedding models
Ollama Ollama /ollama/api/tags, /ollama/api/show, /ollama/api/ps, /ollama/api/version and /ollama/api/pull – the discovery verbs an Ollama tool calls before anything else. The verbs that write to a local model store — create, copy, push, delete — are refused, since this deployment stores no models of its own. The /api/blobs/{digest} upload joined that refused set in v1.18.0, which also answers Ollama is running on the base URL for a client's connection test Amazon Bedrock Amazon Bedrock
OpenAI OpenAI GET /v1/organization/usage/* – tokens, images, characters and seconds per bucket, per model, per endpoint and — where a deployment identifies callers — per user, read from the metrics the gateway itself writes. An endpoint that measures nothing answers empty buckets rather than a 404. Off by default: it needs USAGE_API and CLOUDWATCH_METRICS, and grouping by user needs CLOUDWATCH_METRICS_USER_DIMENSION as well Amazon CloudWatch Amazon CloudWatch
OpenAI OpenAI GET /v1/organization/costs – what AWS bills this deployment for serving those requests, over the same buckets and the same filters. Needs COST_TRACKING on top of the two settings above Amazon CloudWatch Amazon CloudWatch
OpenAI OpenAI POST /v1/realtime/calls – the WebRTC transport: a browser posts its own SDP offer, negotiates a media session directly with the gateway and the audio path terminates there, with no WebSocket relay in between. An offer whose ICE candidates are all private, link-local or mDNS names is refused with invalid_offer; a deployment on a private network sets REALTIME_WEBRTC_ALLOW_PRIVATE_CANDIDATES. Off by default: it needs REALTIME_WEBRTC_ENABLED, and the task must be directly reachable over UDP — behind an HTTP-only load balancer the SDP exchange succeeds but the call carries no audio Amazon Bedrock Amazon Bedrock - Amazon Nova Sonic

Multi-Tenancy

Provider Endpoint/Feature AWS Backend
stdapi.ai Per-tenant API keys – issue a key per tenant, scope it to the models and endpoints that tenant may use, and revoke it without touching anyone else. Files, vector stores, batches, conversations and stored responses stay deployment-wide with no tenant ownership — a documented limitation of this release Amazon DynamoDB Amazon DynamoDB
stdapi.ai Key delivery that skips Terraform state – the server mints each tenant's secret and delivers it once through a parameter, encrypted with the key named by TENANT_KEY_SSM_KMS_KEY_ID — the deployment's own KMS key when the Terraform module provisions it — so reading it takes a grant on that key and not merely permission on the parameter path AWS Systems Manager AWS Systems Manager Parameter Store, AWS KMS
stdapi.ai Tenant AWS credentials – a tenant may bring its own AWS account: its requests then run under its own role, its own Amazon Bedrock quota and its own bill Amazon Bedrock Amazon Bedrock, AWS STS

Model Discovery

Provider Endpoint/Feature AWS Backend
stdapi.ai Wildcard model names – any request that names a model may name a pattern instead — claude-sonnet-*, say — and the most recently released match serves it; an ambiguous tie is refused rather than guessed Amazon Bedrock Amazon Bedrock
stdapi.ai GET /search_models – a model= filter lists everything a pattern matches, newest first, so a pattern can be checked before anything relies on it Amazon Bedrock Amazon Bedrock
stdapi.ai Shared model list – a fleet discovers models once and shares the result, instead of once per task, and the refresh stays off the request path Amazon DynamoDB Amazon DynamoDB

API Improvements

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI Diarized transcripts streamed – a diarized transcription is returned as speaker-labelled segments as they are recognized, instead of being refused Amazon Transcribe Amazon Transcribe
Cohere Cohere Fused multimodal embeddings – text and an image embedded together into one vector, on both embed routes Amazon Bedrock Amazon Bedrock - embedding models
stdapi.ai Cohere Embed v3 listed as image-capable – Bedrock publishes cohere.embed-english-v3 and cohere.embed-multilingual-v3 as text-only although both embed an image and are billed for it; the catalogue now declares the modalities the listing omits, so /v1/models, /search_models and the Models page agree with what is served Amazon Bedrock Amazon Bedrock - embedding models

Platform Features

Feature Description
Public Models page Browse, filter and compare every model the gateway serves, with each fact resolved from the source that states it
Denials name the permission they need An AWS authorization failure is logged with the action that was denied, instead of leaving the operator to infer it from the call that failed
MCP server identity The published container images declare the MCP server they serve, so a registry and an MCP client can identify it
WebRTC media path provisioned Off by default: it needs REALTIME_WEBRTC_ENABLED on the server and realtime_webrtc_media_enabled on the Terraform module. In media mode the module provisions the whole UDP path on the VPC it creates — the matching egress, and the application-subnet network ACL entries the media needs; bring your own subnet_ids and those entries stay yours to widen. The module's realtime_webrtc_stun_server defaults to a Google-operated public STUN server, which the task reaches over an egress rule open to 0.0.0.0/0 on that port — the one place this deployment talks to a third party. Point it at a STUN server you run to keep the flow inside your account
Every setting reachable from Terraform The module carries every setting the server accepts, so an existing table, tenant key prefix or KMS key can be named instead of created, and the grants follow what is named
Per-user attribution role, built by the module aws_bedrock_user_role_create creates the role the gateway assumes per end user — its trust policy and the invoke permissions AWS authorizes against the caller included — instead of leaving it to be written by hand and named by ARN; aws_bedrock_user_role_arn still takes an existing one and wins where both are set. Off by default, as the feature itself is
Scales on CPU out of the box A deployment that sets no autoscaling variable now scales on CPU at 70%, between one task per Availability Zone and five times that — size the ceiling before it surprises the bill
Home Assistant voice, end to end The deployment sample now configures the conversation agent as well as speech, so a spoken command is understood and acted on rather than only transcribed and spoken back

Fixes

  • Images reach the models that read them: the Qwen vision-language line was published by the open-weight family, which advertises text only, so an attached image was refused by models that in fact read it. The line now declares image input and is excluded from the open-weight matcher, so iteration order can no longer decide what a model is published as.
  • The Anthropic MCP connector no longer blocks a request: an mcp_servers request was rejected outright and the mcp_tool_use and mcp_tool_result blocks it produces had no accepted shape, so a client using the connector could not call the API at all. The field and the blocks are accepted, and the request is served without the connector rather than refused — it names an extra tool source, it is not the request itself.
  • An embedded image is billed for what AWS charges: Cohere Embed v3 meters an image as its own billed unit and reports no token count for it, so image embeddings were attributed at zero cost while AWS billed for them. The images the response reports are now recorded against the model's image rate; later Embed versions bill images inside the tokens and are unaffected.
  • A model both catalogues name stays reachable through bedrock-runtime: naming a dual-homed model in AWS_BEDROCK_MANTLE_PREFERRED_MODELS published the preferred entry under that identifier and dropped the other, so the model was reported impossible to batch — batch inference runs on bedrock-runtime alone. The displaced bedrock-runtime entry is kept off-catalogue and read by everything that must reach that endpoint, while the published catalogue keeps naming the preferred entry. It now matters to every deployment, that setting having gained a default.
  • Listings report the time they are ordered by: batch and completion listings paged in identifier order while reporting a creation time recorded from another clock, so the sequence a client saw did not match the timestamps it was given and a page boundary could revisit or skip an entry. The reported time is now read from the identifier itself, and a batch listing finds the newest batches whatever the bucket holds, including after a dense burst.
  • A one-token answer is served rather than refused: the Anthropic Messages API and Chat Completions both accept a budget of a single token — a client probing a model sends exactly that — while the Responses transport a Bedrock Mantle model is served over refuses anything under 16, so such a request came back as a 400 naming max_output_tokens, a field the caller never sent. The budget is now raised to what the transport accepts.
  • An unknown model is a 404 when counting input tokens: POST /v1/responses/input_tokens answered 400 for a model that does not exist, where the official API answers 404 model_not_found — a client branching on the status read a malformed request instead of a missing model. The image routes keep their 400, which is what upstream answers there. Token counting remains unavailable for models served by Bedrock Mantle or a model endpoint, which is a different 400.
  • A deployment on your own IPv4-only subnets stops being told it is dual-stack: the VPC module read "has IPv6" from a subnet attribute AWS reports as an empty string rather than as absent, so every subnet_ids deployment looked dual-stack whatever its subnets carried. The load balancer was then asked for dualstack, an AAAA record was published for an address nothing answered on, and the server was told to bind ::. The module now reads the subnet's actual IPv6 block.
  • A model's release date is when it launched, not when your region got it: a region that opens a model months after launch reports that rollout as the model's own start-of-life, and whichever region happened to be listed first decided the published date — Claude Haiku 4.5 reads as 2026-01-06 in the seven regions it expanded into and 2025-10-15 in the twenty-four that had it at launch. The earliest date any region publishes is now the one reported, by created in the OpenAI model list, by the Anthropic and Ollama model routes, and on the Models page.
  • The Models page states only what a live source backs: a benchmark row is matched to the model release it was measured on, a source that publishes no row now fails the build instead of leaving stale figures in place, facts no longer leak between variants of a family, and a lifecycle date follows one stated rule — the Bedrock API is the reference, and a model card fills a date only where the API states none.
  • A deployment on your own subnets can be planned again: the VPC module this one builds on fed a set to a function that takes only lists, so every plan that creates a VPC failed outright — including for a deployment that had changed nothing, since the constraint resolves to the newest release. Fixed in JGoutin/terraform-aws-vpc v1.6.1.
  • A deployment on your own subnets can be planned again, twice over: the module's IPv6 out-of-band egress rule read whether the subnets carry IPv6 before checking that the WebRTC media path was even on. On subnet_ids that answer comes from the subnets themselves and is unknown while planning, so the plan failed outright — with the rule set empty and WebRTC off, which is every such deployment.
  • A module that names your own VPC no longer deadlocks the plan: the length limit name_prefix carries when a load balancer is enabled was asserted on the variable, which tied it to alb_enabled and through that to subnet_ids. Passing the module's name_prefix output to a VPC and that VPC's subnets back — what the output exists for — closed a dependency cycle. The limit is now checked on the load balancer itself and still refuses too long a prefix while planning.
  • The documentation pages load a current Swagger UI: swagger-ui-dist 5.32.15, which scopes its HTML sanitiser to itself instead of mutating the page's global one. Both pins moved on in v1.18.0, to swagger-ui-dist 5.33.0, which lays out only the part of a large specification that is on screen, and redoc 2.5.4, which carries accessibility fixes and patched bundled dependencies.

v1.16.0 – 2026-08-21 – Conversations, Batches, Vector Stores, Realtime Speech & Per-User Identity (with v1.16.1 maintenance update, 2026-08-25)

At a glance

This release adds four API surfaces and finishes the speech story. New APIs: Conversations keep a thread server-side, so a client continues it by id instead of resending the history; the OpenAI Batch API and Anthropic Message Batches API run large request sets asynchronously at the discounted batch price; Vector Stores index and search files by meaning — or address a knowledge base you already run — with a model reaching either kind for itself through file_search; and the Realtime API holds a spoken conversation over one WebSocket. Speech: 100,000-character synthesis spoken as it is produced, live transcription needing no bucket, and Amazon Nova Sonic as the lowest-cost speech-to-text backend here. Identity per caller: Amazon Cognito tokens alongside or instead of the API key, published discovery so an agent authenticates itself, and per-user cost attribution reporting each end user's spend from the AWS invoice rather than an estimate.

New Required IAM Permissions

v1.16.0 adds one action every deployment needs, a handful that belong to statements you may already grant, and one statement per optional feature. See IAM Permissions for the policies in full.

Enough of them together that they no longer fit one policy: IAM caps a customer managed policy at 6,144 characters, and a deployment enabling most of these exceeds it. Attach several policies to the role rather than widening actions to save room — the Terraform module now ships two, one for Amazon Bedrock and one for the supporting services, and does that for you.

Required on upgrade, whatever the deployment does:

  • bedrock:InvokeModelWithBidirectionalStream — serves every model invoked over a two-way connection — the Realtime API, and Amazon Nova Sonic transcription and translation. It belongs to the core Bedrock policy; no route-specific action exists for any of them.

Add to a statement you already grant, if the deployment uses that feature:

  • bedrock:UpdateSession, on the session storage statement — the conversation metadata update (POST /v1/conversations/{id}) and nothing else. The rest of the Conversations API uses the actions stored responses already require.
  • bedrock-mantle:CountTokens, on the Bedrock Mantle statement — counts the tokens of a Mantle-served model on /anthropic/v1/messages/count_tokens, since Amazon Bedrock's own CountTokens takes Anthropic models only. Needed by any deployment serving Mantle models, which is the default.
  • transcribe:StartStreamTranscription, on the speech-to-text statement — serves stream=true on /v1/audio/transcriptions. Add it with the upgrade: without it a streamed request that names its language answers 503 feature_unavailable, and the server log names the permission. It stages nothing, so a deployment with no bucket at all grants this one alone.
  • polly:StartSpeechSynthesisStream, polly:StartSpeechSynthesisTask and polly:GetSpeechSynthesisTask, plus s3:PutObject, s3:GetObject and s3:DeleteObject on each bucket serving an Amazon Polly Region, on the text-to-speech statement — they serve input above 3,000 characters and nothing else. With a bucket configured and these missing, long requests are accepted and then fail on the permission: grant the whole set, or leave the bucket unconfigured and keep the 3,000-character answer.
  • translate:ListLanguages, on the translation statement — read once at startup so an unsupported language pair is refused before the audio is transcribed. Genuinely optional: without it the check stays off and translation still works, reporting the unsupported pair once the translation call itself fails.

New statements, one per optional feature:

  • Vector stores — s3vectors:CreateIndex, DeleteIndex, PutVectors, GetVectors, QueryVectors and DeleteVectors, scoped to your vector bucket and its indexes. No bucket-level create or delete is granted: the gateway creates and deletes the indexes inside the bucket, never the bucket.
  • Knowledge base vector stores — bedrock:GetKnowledgeBase, Retrieve, ListDataSources, IngestKnowledgeBaseDocuments, ListKnowledgeBaseDocuments, GetKnowledgeBaseDocuments and DeleteKnowledgeBaseDocuments, one statement per allowlisted knowledge base ARN. bedrock:ListKnowledgeBases is deliberately not granted and not needed — the server only ever addresses the identifiers it was given.
  • Batch inference — bedrock:CreateModelInvocationJob, GetModelInvocationJob and StopModelInvocationJob on the server's role, plus iam:PassRole conditioned on bedrock.amazonaws.com. The service role Amazon Bedrock assumes carries its own policy: s3:GetObject, s3:PutObject and s3:ListBucket on the batch prefix, and bedrock:InvokeModel on the models you batch.
  • Per-user cost attribution — sts:AssumeRole and sts:TagSession on the server's role and in the end user role's trust policy (both actions: without TagSession, every tagged session is denied), and bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream and bedrock:ApplyGuardrail on the end user role itself, since AWS authorizes those against the caller of the invocation.
  • Web search — bedrock-websearch:InvokeSearch and bedrock-websearch:InvokeFetch, plus bedrock-websearch:ExternalWebAccess only where a request may reach the open internet. Leaving that last one out is what keeps every search inside the AWS boundary. A missing web-search permission produces no error and no server log entry: the model answers without having searched, so check these before suspecting the model.
  • Transcription output encryption — kms:GenerateDataKey and kms:Decrypt on the key named by AWS_TRANSCRIBE_OUTPUT_ENCRYPTION_KEY_ARN, in the key policy as well as on the role.
  • Durable vector store indexing — sqs:SendMessage, sqs:ReceiveMessage, sqs:DeleteMessage, sqs:ChangeMessageVisibility and sqs:GetQueueAttributes, on the single queue named by AWS_SQS_VECTOR_STORE_QUEUE_URL and never on *. Needed only if you configure that queue; leave it unset and nothing here applies. No queue is ever created, deleted or reconfigured, so none of those actions is granted.

Five features stay inert until you create the resource they need

Nothing in this release is breaking — but these five answer 503, or stay off, until the resource exists in your own account:

  • Batches — an IAM service role Amazon Bedrock assumes to read the requests and write the results (AWS_BEDROCK_BATCH_ROLE_ARN), plus the bucket it reads and writes.
  • Vector stores — an Amazon S3 vector bucket you create yourself, and the Region it lives in.
  • Knowledge base vector stores — an allowlist of the knowledge bases this deployment may address (AWS_BEDROCK_KNOWLEDGE_BASE_IDS), empty by default. One that is not on it answers exactly as a store that does not exist, so the setting cannot be probed for what a deployment holds.
  • Cognito authentication — a user pool and its app clients (AWS_COGNITO_USER_POOL_ID); until then the API key remains the only method, exactly as before.
  • Per-user cost attribution — a role for the end user sessions (AWS_BEDROCK_USER_ROLE_ARN); off by default, and every call keeps being billed to the deployment's own identity until it is set.

Conversations, the Realtime API and streamed transcription need no new resource. Long speech input needs a bucket for the serving region, which is the same one the rest of the gateway already uses — except on generative voices, which speak up to 20,000 characters without one.

Behavior Changes

Review these before upgrading — they may change what existing clients or dashboards observe:

  • A missing deployment permission is no longer reported as the caller's. An AccessDeniedException on the gateway's own AWS calls reached clients as 403 permission_error — which every OpenAI and Anthropic SDK reads as their key being refused. Every route now answers 503 feature_unavailable, with the server log naming the operation, model and permission. Clients matching 403 for a backend permission error should match 503/feature_unavailable instead; a 403 now means only that per-user attribution is on and that end user's role was denied.
  • Built-in web search now appears in usage and cost reporting. Queries were recorded as nothing at all, so a measured turn under-reported its cost by 58%. Nothing AWS charges changed; what the gateway reports does. Web access is also an operator setting now (AWS_BEDROCK_EXTERNAL_WEB_ACCESS), defaulting to the previous behaviour.
  • A request that would be answered without what it asked for is refused. A /v1/responses web_search restricting its sources (filters.allowed_domains, user_location) was accepted and dropped, so answers came back sourced from domains the caller had excluded. Now a 400 on models that cannot serve it; Bedrock Mantle models receive the options unchanged. The same rule governs file search filters and score thresholds.
  • Two output-shaping hints that returned 400 now succeed. prediction and verbosity on chat completions are accepted and dropped, as the Responses surface already did; truncation="disabled" is likewise accepted, while truncation="auto" is still refused.
  • /v1/responses forwards undeclared request fields to the model, as chat completions and messages already did, so the backend may refuse one it does not recognise. Conversely, client-side control fields no provider treats as parameters (LiteLLM's drop_params among them) are dropped rather than forwarded. Both are governed by EXTRA_MODEL_PARAMS_DENYLIST and EXTRA_MODEL_PARAMS_DROP_ALL.
  • Attachments are measured against what the model actually accepts. The old guard compared raw bytes where the backend enforces base64 length, so it was ~33% too permissive. Oversized attachments are now staged and referenced where the model reads from storage, or refused with 413 naming the size it accepts. Smaller attachments are unaffected — see Attachment Size.
  • The server's own connections follow the proxy environment. HTTPS_PROXY, HTTP_PROXY and NO_PROXY were honoured by the AWS SDK and ignored by everything else, so a proxied deployment saw no Bedrock Mantle models. Two connections deliberately still bypass it: container metadata, and the fetch of a caller-supplied URL, where a proxy would defeat address validation. See proxied deployments.
  • A declared upload checksum is now verified. The value was stored and never looked at, so a corrupted upload completed like a clean one. It covers the file's contents, not the storage layer's multipart identifier — declaring the latter is now refused.
  • An unknown model name answers with a sentence, not the catalogue. The 404 body carried every served identifier, roughly 2,500 characters. Clients that parsed it for a model list should call /v1/models.
  • Bedrock Mantle is only probed in the Regions that serve it, so a deployment listing others no longer warns at every start. An explicit AWS_BEDROCK_MANTLE_REGIONS list is still used exactly as given.
  • The container health probe's command changed. Deployments that re-declare the probe instead of running the image's own — an ECS task definition among them — should take the command from the image.

New APIs

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/conversations – create, retrieve, update and delete a conversation, list and manage its items, and continue it from the Responses API with the conversation parameter Amazon Bedrock Amazon Bedrock - session management
OpenAI OpenAI /v1/vector_stores – attach files, follow the indexing as it progresses, then search by meaning with attribute filters and per-passage scores Amazon S3 Amazon S3 Vectors, Amazon Bedrock Amazon Bedrock - embedding models
OpenAI OpenAI /v1/vector_stores – address an Amazon Bedrock knowledge base you already run as a vector store, Bedrock managed or customer-managed: search it, attach, list, read and delete documents. Allowlisted per knowledge base, never created or deleted here Amazon Bedrock Amazon Bedrock - Knowledge Bases
OpenAI OpenAI file_search on /v1/responses – a chat model answers from the stores you name, reporting the searches it ran and citing a file_citation per file it drew on Amazon S3 Amazon S3 Vectors, Amazon Bedrock Amazon Bedrock - Knowledge Bases
OpenAI OpenAI /v1/batches – run a JSONL file of chat completion or embedding requests asynchronously at the batch price: submit, poll, cancel, read the result files Amazon Bedrock Amazon Bedrock - batch inference
Anthropic Anthropic /anthropic/v1/messages/batches – the same asynchronous, batch-priced run for the Messages API, results streamed back as JSONL; each request may name its own model, up to eight per batch Amazon Bedrock Amazon Bedrock - batch inference
OpenAI OpenAI WS /v1/realtime – a live speech-to-speech session over one WebSocket, with a transcript of both sides, server-side turn detection or manual turns, barge-in, and G.711 for telephony Amazon Bedrock Amazon Bedrock - Amazon Nova Sonic
OpenAI OpenAI POST /v1/realtime/client_secrets – mint a short-lived, browser-safe credential carrying a session configuration; signed and stateless, so any instance verifies one minted by any other Amazon Bedrock Amazon Bedrock - Amazon Nova Sonic

Limits worth knowing before building on these

Realtime: a session lasts at most 8 minutes and calls no tools, and a spoken answer is guardrail-checked once complete, so a blocked one may already have been partly heard (coverage). WebRTC and SIP are not served in this release — put LiveKit Agents or Pipecat in front for a browser media path or a phone line. Superseded in v1.17.0, which added the gateway-terminated WebRTC transport as an operator opt-in; SIP is still never terminated here. Tool calling was added in v1.18.0, so a session does call tools. The compatibility table lists every event the session does not emit.

Knowledge-base stores address a knowledge base that already exists and refuse, naming why, anything that would reshape it — creating, deleting, renaming, expiry, chunking strategy, attribute rewrites and the file-batch routes. Attaching needs a custom data source. Retrieval scores are reported as the backend states them rather than rescaled into similarities, and unknown values are reported unknown rather than invented. See Knowledge Base Stores.

Speech & Audio

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/audio/speech – up to 100,000 billed characters per request, 24× the upstream 4,096, with no API change and no new request field Amazon Polly Amazon Polly
OpenAI OpenAI /v1/audio/speech – long input is spoken as it is synthesized instead of after a whole job finishes; generative voices reach 20,000 characters with no bucket at all, and each request takes whichever path can serve it, so long input is no longer tied to one voice Amazon Polly Amazon Polly
OpenAI OpenAI /v1/audio/transcriptions – naming Amazon Nova Sonic transcribes at the lowest cost available here, streamed as it is recognized; json and text only, up to 10 minutes, no timestamps. No existing request is re-routed Amazon Bedrock Amazon Bedrock - Amazon Nova Sonic
OpenAI OpenAI /v1/audio/translations – Amazon Nova Sonic translates speech to English itself, in one request Amazon Bedrock Amazon Bedrock - Amazon Nova Sonic
OpenAI OpenAI /v1/audio/transcriptions – stream=true returns each phrase as it is recognized, whenever the request names the language to expect; needs no bucket. gpt-live-transcribe is now an alias, and requests naming no language are unchanged unless AWS_TRANSCRIBE_STREAM_LANGUAGES says which to expect Amazon Transcribe Amazon Transcribe
OpenAI OpenAI /v1/audio/transcriptions – per-language custom vocabularies and language models, so a request identifying between several languages can apply the right resources to each one instead of being refused. Accepted only where the backend would use them: alongside a single fixed language, where they would apply to nothing, they are still refused Amazon Transcribe Amazon Transcribe
stdapi.ai AWS_TRANSCRIBE_OUTPUT_ENCRYPTION_KEY_ARN – encrypt a transcription's output with a key you name rather than the bucket's own. The job's request identifiers travel as the encryption context, so a key policy can be scoped to this workload instead of to the whole bucket Amazon Transcribe Amazon Transcribe, AWS KMS
OpenAI OpenAI /v1/audio/translations – the supported language pairs are read once at startup and checked before the call, so a pair that cannot be served is named as the request problem it is instead of surfacing as a failure after the audio was transcribed. The permission that reads them is optional: without it the check stays off and everything else works AWS Translate AWS Translate

Identity & Cost Attribution

Provider Endpoint/Feature AWS Backend
stdapi.ai Amazon Cognito user pool tokens – accept access tokens instead of, or alongside, the API key, so each caller reaches the API with their own credential; validated in-process against the pool's published keys, with no AWS call on the request path Amazon Cognito Amazon Cognito
stdapi.ai AUTHENTICATION_MODE – assert the posture rather than infer it: the server refuses to start when the selected method is not configured, or when a configured method would be silently ignored Amazon Cognito Amazon Cognito
stdapi.ai Authentication discovery for agents – an OAuth 2.0 protected resource metadata document, pointed at by every unauthorized response, so an MCP client finds the authorization server and the scope it needs without being configured for this deployment; published only once an authorization server is declared Amazon Cognito Amazon Cognito, or any OAuth 2.0 authorization server
stdapi.ai Per-user cost attribution – model calls issued under a short-lived role session tagged with the caller, so AWS reports each end user's spend in Cost Explorer and the Cost and Usage Report, from the invoice rather than an estimate. Off by default; a deployment can also require every call to name its end user rather than bill it to the deployment Amazon Bedrock Amazon Bedrock, AWS STS
stdapi.ai Vector store cost reporting – a search against a Bedrock-managed knowledge base is recorded and priced like every other billed unit; what cannot be accounted for is stated rather than approximated Amazon Bedrock Amazon Bedrock - Knowledge Bases

Platform Features

Feature Description
Aliases that carry configuration A MODEL_ALIASES entry may map a public name to the target model plus the service tier, guardrail, metadata and extra parameters applied to requests naming it, so one model is published under several names with different policies. The plain-string form is unchanged, and a malformed alias stops startup naming itself rather than failing once per request
Attachment size policy On the multimodal routes served by Amazon Bedrock, an attachment is measured before the request is built and travels inline or by reference according to the limits each model class declares. Staging is per model and per media kind: of the families measured, only the Amazon Nova families and TwelveLabs Pegasus accept a reference
Ephemeral secret signing key REALTIME_CLIENT_SECRET_KEY signs the Realtime API's client secrets. A deployment with an API key already shares one and needs nothing; one running with no API key at all should set it, or a secret minted by one instance fails to verify on another
Dual-stack container listener The image's bind address moved out of its command into GRANIAN_HOST, so a deployment that needs a dual-stack socket — an ECS service whose discovery record includes an AAAA record, for instance — sets one variable instead of replacing the whole command. The IPv4-only default is unchanged
Faster container health probe The probe ships as a module of the application itself, byte-compiled with the rest of the package and covered by the linters and the test suite; it speaks HTTP over a socket rather than pulling in 123 modules per probe, cutting roughly 250 ms of import work per run in the community image and halving its peak memory
OpenAI Daybreak models Daybreak Red (GPT-5.6 Cyber) and Daybreak Blue (GPT-5.6 Sol) are served and priced with the rest of the GPT-5.6 family, image input included. Both answer on the Responses API through Bedrock's next-generation inference endpoint, in US East (Ohio) only, and both are gated on enrollment with OpenAI's Daybreak programme — an account without it does not see them in the catalogue at all
Capability discovery The model catalogue advertises what this release added — speech to speech, its transcription and translation, the search surfaces, and whether a model can be used with the Batch API — filterable over HTTP and through the same tool an agent reads before it calls anything. Web search is credited to every model that provides it, not only to the family the last release added it for
Durable vector store indexing AWS_SQS_VECTOR_STORE_QUEUE_URL hands indexing to an Amazon SQS queue you create, so a file keeps being indexed — and finishes — when the server that accepted it is replaced. Off by default; needs a standard queue with a dead-letter queue and the durable indexing permissions (resilience)
End-to-end client coverage The suite driving complete, unmodified third-party clients against a live gateway gains five: LiteLLM, Docling Serve's vision pipeline, the OpenAI Agents SDK (realtime voice, conversations, web search, vector-store retrieval), and LiveKit Agents and Pipecat running the exact WebRTC and telephony configurations this documentation prints. Deployment guides follow for LobeHub and RAGFlow

Fixes

  • Vector store durability: an attached file whose indexing was interrupted is now reported failed with a last_error instead of sitting in_progress for ever; a detached file leaves every listing and search immediately and is gone only once its passages are, so a server replaced mid-delete no longer leaves a deleted document answering searches; indexing is bounded for the whole server rather than per request, so memory and embedding quota no longer scale with the number of callers; and deployments that would rather the work finished than reported can hand indexing to a queue
  • Work a request left running is finished before the server stops: a deployment, scale-in or Spot interruption used to drop temporary file cleanups, vector store indexing and live audio session releases with nothing in the logs. Shutdown now waits under SHUTDOWN_DRAIN_TIMEOUT (10 seconds by default), settles whatever the deadline leaves, and counts it in the stop event. Raise it together with your container runtime's kill delay, never one alone
  • Streamed responses run the work they scheduled: the drain was attached before the body produced a byte, so three leaks followed — a vector store searched only through streamed answers could expire mid-query, an expired store left a paid index behind, and a streamed transcription falling back to a job left its audio, transcript and job record behind on every request
  • Realtime speaks the released vocabulary, not the beta one: item events were unparsable, the caller's transcript event was dropped for a missing field, and the item lifecycle clients wait on was never emitted. Barge-in did not work at all — the session refused truncation precisely while an answer was playing, the only moment it is ever sent. Truncate, retrieve and delete now work against the tracked conversation, a written turn is answered instead of timing out, and every answer reports the six response fields upstream always sends
  • Batch API: listing a batch neither settled nor published it, so a client that only ever listed never had its usage recorded; every validation failure at submission was reported as an unsupported model, quota and role failures included; and the batch record was written only after the jobs started, leaving billable work running with nothing on disk to stop it
  • Knowledge-base stores answer in this API's own words: attaching to a managed knowledge base always failed, listing its files always failed, deleted documents never left a listing, and internal bookkeeping reached the caller as attributes. A file refused because the corpus is maintained elsewhere is now told so in the store's terms, and a refused file is explained by the store that refused it rather than by one fixed sentence describing limits the caller never met
  • Cost reporting matches what AWS charges: GPT-5.6 Luna was reported at five times its real cost and Terra a quarter over, since AWS repriced them on the model's own page rather than in the live price catalogue; a Global cross-Region call — the only way Luna, Sol and Terra are ever served — was costed at the In-Region rate, about 10% over; and moderation usage naming an alias matched no price at all. Every source page is now re-checked weekly
  • Reasoning and web search reach the models that serve them: Amazon Nova 2 and DeepSeek V3 refused the token budget the Anthropic dialect requires, leaving extended thinking unreachable on that route while the identical ask worked elsewhere; and a web_search sent to a GPT-5.6 model resolved to its non-Mantle twin travelled as an ordinary function tool, so no search ran and nothing said so — now a 400 naming both ways to route the model to the endpoint that serves it
  • The Messages surface reports what an answer cost and why it stopped: refusals carry the policy category, the reasoning-token breakdown is reported, and service tier and per-TTL cache-creation counts are populated — most visibly on a batch, which claimed no tier while being billed as one
  • Audio, embeddings and attachments are bounded correctly: inline audio was measured against raw bytes where the backend enforces the encoded length, so files between ~18.75 MB and 25 MB passed and were refused downstream; long text-to-speech now answers with the length a bucket-less deployment can honour; and models that embed one input per call no longer open a connection per chunk
  • The interactive documentation pages render with no outbound access: /docs and /redoc pulled the icon, Swagger UI, ReDoc and a web font from three third parties — blank pages in an air-gapped VPC, and elsewhere a report to those hosts of who was reading this API and when, running whatever a floating major version tag resolved to that day. Both are now served whole from the image, pinned to exact releases verified by SHA-256 at build time, with upstream licences beside them
  • MCP tools return what they produce: every route answering with bytes was published as a tool an agent could call and then could not use. An image now arrives as an image and audio as audio, anything the protocol cannot carry arrives as a reference rather than failing, and the 4 MiB body cap that blocked image edits now follows MAX_INPUT_FILE_SIZE
  • Addresses and listings are the ones this deployment serves: a custom route prefix still quoted default paths to search_models and to video job polling, neither recoverable client-side; and a file's created_at and its place in a listing came from two different clocks, so a multipart upload sat among older files reporting a later time. The Anthropic listing also answered oldest first, hiding every recent file, and now runs newest first as upstream does
  • Diagnostics name their cause: an unreachable Region rendered six different conditions as one identical sentence, and the slow startup beside it was the container metadata lookup retrying, unreported; a request abandoned mid-flight left its OpenTelemetry trace current, so later work was recorded under a closed request's trace id; and behind a proxy, client_ip recorded the load balancer whatever PROXY_TRUSTED_HOSTS allowed
  • The API describes itself, not the service behind it: a synthesis limit credited to the service enforcing it, prices credited to their catalogue and a moderation route naming the engine underneath all shipped in the OpenAPI document and in the tool descriptions agents read before calling

Fixes & Maintenance (v1.16.1)

  • Update cryptography to 50.0.1, rebuilt against OpenSSL 4.0.2

v1.15.0 – 2026-08-03 – Reliability, Performance & Feature Completeness

At a glance

  • The largest correctness pass to date — three successive deep audits plus an independent full-branch review closed hundreds of fidelity gaps across all three API dialects, every fix pinned by tests.
  • Performance — hot paths run in compiled native code and independent work runs in parallel, cutting the gateway's processing overhead.
  • Explicit prompt caching — placement and lifetime control on the OpenAI dialect, plus fixes to the caching plumbing that already existed.
  • Reasoning and prompt controls — operator reasoning controls, the Responses API prompt parameter from Amazon Bedrock Prompt Management, and native mid-conversation system messages on Claude 4.8+.
  • Verified with real clients — compatibility claims are backed by a test tier running complete, unmodified third-party client software against a live gateway. Behavior changes: unsupported parameters are now accepted and ignored, and a configured guardrail applies to every route.

This release focuses on making the whole gateway better rather than just bigger. Reliability and quality: the largest correctness pass to date — three successive deep audits plus an independent full-branch review closed hundreds of fidelity gaps across all three API dialects, every fix pinned by tests and the whole surface validated by real, unmodified client applications. Performance: hot paths now run in compiled native code and independent work in parallel, measurably cutting the gateway's processing overhead. Feature completeness: existing capabilities are rounded out end to end — explicit prompt caching, operator reasoning controls, the Responses API prompt parameter from Amazon Bedrock Prompt Management, native mid-conversation system messages on Claude 4.8+, richer speech and transcription (Polly speech marks, Transcribe/Translate extras, generic Converse speech-to-text), Cohere embedding_types and Rerank v1 structured documents, guardrail enforcement on every route, and an inline guardrail-checks moderation backend.

Behavior Changes

Review these before upgrading — they may change what existing clients observe:

  • Unsupported parameters are accepted and ignored, not rejected. Parameters the AWS backends cannot honour (e.g. known_speaker_*, partial_images, unsupported image quality/style, programmatic tool calling) now behave like they do on OpenAI: the request succeeds, the parameter is dropped, and a warning is recorded in the request log. Requests that returned 400 on v1.14 may now succeed.
  • Impossible combinations are now clean 400s instead of silent degradation: subtitle or diarized formats with stream=true, contradictory Amazon Transcribe settings, and web-search filters Nova grounding cannot apply are rejected with actionable messages.
  • Speech output quality: wav/flac/aac are now encoded from lossless PCM instead of Ogg Vorbis, and the default pcm output is resampled to 24 kHz for OpenAI parity (pass an explicit SampleRate to keep Polly's native rate). Same formats, different — better — bytes.
  • Error responses no longer expose backend internals. Server-side (5xx) error messages are generic with details kept in the server log, and Anthropic error types now match the official SDK exactly.
  • Comprehend-backed moderation always analyses text as English — the only language the AWS API accepts at runtime.
  • A configured guardrail now applies to every route. Embeddings, rerank, images, videos, and the audio routes enforce it through the ApplyGuardrail API (route coverage) — requests that silently bypassed the guardrail on v1.14 may now return 400 (code content_filter) or masked text, and each check is billed as guardrail text units.
  • SSRF protection covers every non-globally-reachable address. With SSRF_PROTECTION_BLOCK_PRIVATE_NETWORKS enabled (the default), a user-supplied URL resolving to shared address space (100.64.0.0/10, used by EKS custom networking and Hybrid Nodes) or another special-purpose range is now rejected with 403, alongside the RFC 1918 ranges.
  • Usage reporting is additive but richer: cached-token buckets are folded into prompt_tokens with prompt_tokens_details on every surface, and the Anthropic API now reports cache_creation_input_tokens (it was always null on v1.14).
  • Two request-body keys are reserved. model_id and additional_request_fields (plus stop_sequences on the legacy /v1/completions, where stop is the parameter to use) collide with the gateway's own request-building parameters: instead of being forwarded to Bedrock as provider extras, they return a 400 invalid_request_error naming the key.
  • The container health probe now respects TRUSTED_HOSTS. The image's HEALTHCHECK requests /health with a Host header derived from TRUSTED_HOSTS — a correct list keeps the container healthy with no extra entry. Deployments that re-declare the probe, such as an ECS task definition, should run the image's own command. Note that a load balancer health check still sends the target's IP address as the Host and is rejected with 400 when the allow-list is enabled.

Explicit Prompt Caching

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI Explicit cache breakpoints on chat completions and responses, mapped to Bedrock cachePoint blocks with the per-request block budget enforced Amazon Bedrock Amazon Bedrock - Converse API
OpenAI OpenAI prompt_cache_options – cache TTL control on chat completions and responses, honoured on models that support it Amazon Bedrock Amazon Bedrock - Converse API

Prompt caching itself is not new — this release adds explicit placement and lifetime control on the OpenAI dialect, and fixes the caching plumbing that already existed (see Fixes).

Reasoning Controls

Provider Endpoint/Feature AWS Backend
stdapi.ai CHAT_COMPLETIONS_REASONING_FIELD – return thinking text under reasoning_content, reasoning, or suppress it with none; applied identically to streamed deltas and final messages Amazon Bedrock Amazon Bedrock - Converse API & Mantle
OpenAI OpenAI OpenRouter-style reasoning request object accepted on chat completions (effort, max_tokens, enabled, exclude) Amazon Bedrock Amazon Bedrock - Converse API & Mantle

New API Features

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI Responses API prompt parameter – serve prompts stored in Bedrock Prompt Management, with versions and variables (opt-in, AWS_BEDROCK_ALLOW_PROMPT_ARN) Amazon Bedrock Amazon Bedrock - Prompt Management
Anthropic Anthropic / OpenAI OpenAI Mid-conversation system messages forwarded natively on Claude 4.8+ and Claude 5 family models instead of being folded into the system prompt Amazon Bedrock Amazon Bedrock - Converse API & Mantle
stdapi.ai Guardrail asynchronous stream processing via the X-Amzn-Bedrock-GuardrailStreamProcessingMode request header Amazon Bedrock Amazon Bedrock - Guardrails
stdapi.ai Configured guardrails enforced on every route: embeddings, rerank, images, videos, and audio now apply them via the ApplyGuardrail API (route coverage) Amazon Bedrock Amazon Bedrock - Guardrails
OpenAI OpenAI Moderations amazon.bedrock-runtime-guardrail-checks model – inline guardrail content filter checks with no guardrail resource required, the new default fallback for omni-moderation-* in supported regions Amazon Bedrock Amazon Bedrock - Guardrails

Speech & Audio

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/audio/speech – Polly SpeechMarkTypes for word/sentence/viseme/SSML timing marks Amazon Polly Amazon Polly
OpenAI OpenAI /v1/audio/transcriptions – Amazon Transcribe extra parameters: multi-language identification, custom vocabularies, PII redaction, and more Amazon Transcribe Amazon Transcribe
OpenAI OpenAI /v1/audio/transcriptions – gpt-transcribe context inputs: keywords and multi-language languages, with detected languages reported in the response Amazon Transcribe Amazon Transcribe
OpenAI OpenAI /v1/audio/transcriptions & translations – any Converse-capable speech-input Bedrock model transcribes through a generic default (Voxtral rebuilt on the Converse API); uploads outside the accepted formats are transcoded automatically Amazon Bedrock Amazon Bedrock - Converse API
OpenAI OpenAI /v1/audio/translations – AWS Translate Formality, Profanity, Brevity, and custom terminologies Amazon AWS Translate

Cohere Embed & Rerank

Provider Endpoint/Feature AWS Backend
Cohere Cohere embedding_types – quantized (int8/uint8/binary/ubinary) and base64 embeddings on both embed routes; image embedding metadata reported Amazon Bedrock Amazon Bedrock - embedding models
Cohere Cohere Rerank v1 – structured JSON documents with rank_fields selection Amazon Bedrock Amazon Bedrock - Rerank API

Platform Features

Feature Description
Stateless MCP transport The MCP server can serve /mcp without server-side sessions (MCP_STATELESS_HTTP), so any replica may answer any request, alongside a GET /ping health probe kept out of request logs
Retry-After on 429 Throttled responses advertise the region router's computed backoff, so well-behaved clients retry exactly when capacity returns
AWS request-ID correlation Request logs record every AWS API call's request ID (and incoming ALB/CloudFront trace headers), so a gateway request ties directly to CloudTrail and AWS support cases
Programmatic tool calling types The OpenAI SDK's programmatic tool calling type surface parses on every request union, accepted and ignored on models without the capability
Performance Hot paths run in compiled native code and independent work runs in parallel: a 1 MB request costs 30% less CPU, multi-image generations finish in the time of the slowest image, and every optimization is pinned by regression tests
MCP context efficiency Tool schemas hide parameters MCP callers cannot use (streaming modes, token-level tuning, caller identifiers) and tool results return as compact JSON, cutting the tokens each call costs the calling agent; every exposed MCP tool is exercised end to end through a real MCP client in the test suite
Slimmer container image The community image shrinks from 230 MB to 156 MB: unused dependency payloads are removed (language-name data, cryptography, rich/typer, uvicorn extras) and AWS service models are pruned to the services actually used, guarded by a build-time smoke test; the package inventory stays complete for vulnerability scanners (self-built ffmpeg registered, Python package metadata retained) and every redistributed component keeps its licence and notice files
Metadata filter for MCP clients Listing stored chat completions accepts the metadata filter as a single metadata={"key": "value"} JSON object as well as the OpenAI SDK's metadata[key]=value pairs, so clients that can only send one query parameter per field — MCP tool calls among them — can filter too

Verified with Real Clients

Compatibility claims in this release are backed by a new test tier that runs complete, unmodified third-party client software against a live gateway — not just HTTP assertions. Coding agents (Claude Code, Codex, pi, Qwen Code), the n8n workflow platform, Open WebUI, Home Assistant's voice bridge, a Haystack RAG pipeline, and the LangChain and pydantic-ai libraries drive real multi-turn tool-calling, retrieval, and speech sessions across dozens of models and all three API dialects, in isolated sandboxes. Alongside them, every served model is empirically probed for the parameters it genuinely honours, with the results recorded and pinned by tests. See Quality Assurance for the full methodology.

Fixes

Three audit passes and an independent full-branch review closed over a hundred fidelity gaps. The user-visible highlights:

  • MCP tool calls match their schemas: union-typed parameters no longer advertise a contradictory single type (which made valid string arguments randomly fail schema validation), and the JSON image edit/variation bodies accept the plain string references the tool schemas advertise
  • Clean errors on log-exempt paths: a 404 or 405 on paths kept out of request logs (e.g. /favicon.ico, auto-requested by every browser visiting /) returned a 500 with a traceback instead of the JSON error envelope
  • Prompt caching plumbing: Anthropic cache_control breakpoints land on their marked block; cache reads and writes are counted, priced, and reported consistently in responses, request logs, and count_tokens
  • Reasoning: reasoning_effort="max" accepted end to end; thinking text returned by Bedrock Mantle models is surfaced on every API instead of dropped; unsigned reasoning is no longer replayed to models that would reject it
  • Streaming parity: streamed and non-streamed results are now identical (tool-call indices, same-role message merging, text concatenation); mid-stream errors emit proper error events on every API instead of ending streams silently; redacted thinking and web-search results round-trip exactly as native Anthropic emits them; failed generations return 502 instead of an empty 200
  • Routing & billing: a read timeout on an already-sent request is no longer re-invoked in another region (no double billing), on Converse and Mantle alike; region failover now tries each candidate region at most once per request instead of looping back over regions it just marked as throttled, so a single request can no longer escalate a region's quota backoff toward the one-hour ceiling; Bedrock prompt-router usage is billed against the actually-invoked model; the price card reprices a standard-tier row served as a tier fallback at the rate that tier actually bills; store=true degrades gracefully with a logged warning in regions without the session API
  • Responses API parity: the type surface is synchronized with the current OpenAI SDK (tool fields, error codes, tool-call caller provenance); Anthropic count_tokens counts exactly what generation sends, and error bodies carry the request_id
  • Audio & images: the transcoding pipeline is fully bounded — a stalled or failed encode returns a clean error instead of holding the connection open; multipart forms bind every list field the OpenAI SDK sends; size="auto" works on generation, edits, and variations; zh-TW/pt-PT stay distinct in translation; PII-redacted transcripts are read from the key Amazon Transcribe actually writes; Polly voice auto-selection is deterministic; batch-purpose files apply the documented 30-day default expiry, and an expired file now disappears from file listings instead of being listed with an entry that 404s on retrieve

v1.14.0 – 2026-07-12 – Bedrock Mantle, Video Generation, Cohere APIs, Moderation & Stored Conversations

At a glance

  • Amazon Bedrock Mantle, enabled by default — the models served by the Mantle endpoint (OpenAI GPT-5.4/5.5/5.6, xAI Grok 4.3, Google Gemma 4, Qwen3, GLM, DeepSeek, MiniMax, Kimi, Nemotron and more) become available through all four text APIs, with independent throughput quotas.
  • A Cohere-compatible API — Rerank and Embed, making this a three-dialect gateway.
  • Videos and moderation — the OpenAI-compatible Videos API for asynchronous video generation, and content moderation backed by Amazon Bedrock Guardrails or Amazon Comprehend toxicity detection.
  • Stored conversations — store=true, previous_response_id continuation and a full lifecycle on Amazon Bedrock session storage, plus conversation compaction and Responses API extended reasoning.
  • Operations — a model pricing API, multi-region failover for every AWS AI service, fault-tolerant startup, real AWS-billed usage and costs in request logs, and a security hardening pass. Two new IAM permissions are required.

This release adds enabled-by-default Amazon Bedrock Mantle support — models served by the Bedrock Mantle endpoint (OpenAI GPT-5.4/5.5/5.6, xAI Grok 4.3, Google Gemma 4, Qwen3, GLM, DeepSeek, MiniMax, Kimi, Nemotron, and more) become available through all four text APIs, with transparent API conversion, native stored conversations, and independent throughput quotas. It also turns stdapi.ai into a three-dialect gateway with the new Cohere-compatible API (Rerank and Embed), adds the OpenAI-compatible Videos API for asynchronous video generation, content moderation backed by Amazon Bedrock Guardrails or Amazon Comprehend toxicity detection, stored responses and chat completions with store=true, previous_response_id multi-turn continuation, and a full list/retrieve/update/delete lifecycle on Amazon Bedrock session storage, and conversation compaction. The Responses API gains extended reasoning: Bedrock reasoningContent now surfaces as native reasoning output items, both non-streaming and streamed, with signatures and redacted payloads round-tripping through an encrypted_content envelope. A broader compatibility pass brings request/response parity closer to the OpenAI SDK — hosted and agent tool types (web search, computer use, custom tools) are now accepted and ignored instead of rejected, streams correctly terminate with response.incomplete/response.failed, cached tokens are counted in input_tokens, and citation annotations are emitted with their streaming events — validated end-to-end against the OpenAI Codex CLI as an agent client. Operations gain a model pricing API, multi-region failover for every AWS AI service, fault-tolerant startup, real AWS-billed usage and costs in request logs (optionally exported as CloudWatch metrics), and a security hardening pass covering SSRF protection, input validation, and log/error redaction.

New Required IAM Permissions

v1.14.0 requires two new IAM permissions:

  • bedrock:Rerank — needed for the Cohere-compatible Rerank API (/cohere/v2/rerank). See IAM Permissions.
  • bedrock:ListAsyncInvokes, plus bedrock:ListTagsForResource on arn:aws:bedrock:*:*:async-invoke/* — needed for GET /v1/videos (listing video generation jobs across regions). See IAM Permissions.

Ensure your IAM role or user policy includes both statements before upgrading to v1.14.0.

Session storage and Comprehend permissions already covered

The IAM permissions for stored responses/chat completions (bedrock:CreateSession and related session actions) and Comprehend-based moderation (comprehend:DetectToxicContent) were already added to the official stdapi-ai Terraform module ahead of this release. Deployments using a hand-written policy still need to add those statements if they haven't already. Without the session permissions, store=true (previously accepted and ignored) is still ignored — a warning is recorded in the request log instead of failing the request.

Amazon Bedrock Mantle

Enabled-by-default support (AWS_BEDROCK_MANTLE_ENABLED) for models served by the Amazon Bedrock Mantle endpoint — OpenAI GPT-5.4/5.5/5.6 (Sol, Terra, Luna), xAI Grok 4.3, Google Gemma 4, Qwen3, GLM 4.x/5, DeepSeek V3.x, MiniMax M2.x, Kimi K2.5, Nemotron, and more — alongside the classic Bedrock Converse catalog:

  • All four text APIs (chat completions, responses, messages, legacy completions) are served for every Mantle model — native passthrough where the model supports the API upstream, transparent conversion otherwise
  • Models available on both bedrock-runtime and Mantle are served by bedrock-runtime by default; AWS_BEDROCK_MANTLE_PREFERRED_MODELS or the opt-in x-stdapi-service request header (AWS_BEDROCK_MANTLE_SERVICE_HEADER) route them through Mantle instead — e.g. to tap Mantle's independent throughput quotas
  • Native Mantle stored conversations on /v1/responses (store, previous_response_id, retrieval and deletion) — 30-day retention, region-local, project-scoped
  • Multi-region failover and quota backoff across AWS_BEDROCK_MANTLE_REGIONS, matching classic Bedrock region routing
  • Authentication via short-term bearer tokens derived from the server's AWS credential chain — no static secrets
  • Usage recorded and priced at bedrock-mantle rates, including cached tokens and service tiers
  • Optional Bedrock Project/Workspace attribution for cost tracking via AWS_BEDROCK_MANTLE_PROJECT, with per-request override (AWS_BEDROCK_ALLOW_MANTLE_PROJECT_OVERRIDE) through the OpenAI-Project / anthropic-workspace header

Bedrock Mantle Models

Additional IAM Permissions (opt-in feature)

Enabling AWS_BEDROCK_MANTLE_ENABLED requires the bedrock-mantle:CreateInference, bedrock-mantle:GetInference, bedrock-mantle:DeleteInference, bedrock-mantle:ListModels, bedrock-mantle:GetModel, and bedrock-mantle:CancelInference permissions on arn:aws:bedrock-mantle:*:*:project/*, plus bedrock-mantle:CallWithBearerToken on *. See IAM Permissions.

New APIs

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/videos – create, poll, list, download, and delete video generation jobs Amazon Bedrock Amazon Bedrock - Amazon Nova Reel, Luma Ray 2
OpenAI OpenAI /v1/moderations – text and image content classification Amazon Bedrock Amazon Bedrock - Guardrails, Amazon Comprehend
Cohere Cohere /cohere/v2/rerank – document reranking (Amazon Rerank 1.0, Cohere Rerank 3.5) Amazon Bedrock Amazon Bedrock - Rerank API
Cohere Cohere /cohere/v2/embed – embeddings over all Bedrock embedding models Amazon Bedrock Amazon Bedrock - embedding models
stdapi.ai /model_pricing – exact AWS unit prices per model AWS Price List API

Extended Reasoning

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/responses – Bedrock reasoningContent returned as reasoning output items Amazon Bedrock Amazon Bedrock - Converse API
OpenAI OpenAI Streaming response.output_item.added / response.reasoning_text.delta / .done events for reasoning content Amazon Bedrock Amazon Bedrock - Converse API
OpenAI OpenAI include=["reasoning.encrypted_content"] – signature/redacted round-trip for multi-turn reasoning continuation Amazon Bedrock Amazon Bedrock - Converse API

Conversations

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI store=true + GET/DELETE /v1/responses/{id}, input items listing, and previous_response_id continuation Amazon Bedrock Amazon Bedrock - session management
OpenAI OpenAI POST /v1/responses/{id}/cancel – endpoint parity for the cancel lifecycle (always fails for session-stored responses, which never run in background mode; Mantle-stored responses are cancelled upstream) Amazon Bedrock Amazon Bedrock - session management
OpenAI OpenAI store=true + GET/DELETE /v1/chat/completions/{id}, GET /v1/chat/completions listing, POST /v1/chat/completions/{id} metadata updates, and input messages listing Amazon Bedrock Amazon Bedrock - session management
OpenAI OpenAI /v1/responses/compact – stateless conversation compaction Amazon Bedrock Amazon Bedrock - Converse API
OpenAI OpenAI moderation request parameter on chat completions and responses, with results reported in the response Amazon Bedrock Amazon Bedrock - Guardrails

Platform Features

Feature Description
Claude 5 models Explicit support for the Claude 5 generation — Opus 5, Sonnet 5, Fable 5, and Mythos — with each model's server tool set and reasoning configuration matched to what Bedrock actually accepts (Opus 5 exposes no computer use tool; Fable and Mythos always reason and reject a disabled configuration). Model matching covers unreleased versions of each family, so a new minor or major release inherits its family's behavior instead of a generic fallback. Validated end-to-end across the full Claude feature matrix, from Claude 4.5 through Claude 5
Multi-region AWS AI services Automatic multi-region failover for Amazon Polly, Transcribe, Translate, and Comprehend (per-engine voice discovery, co-located Transcribe buckets, latency-ordered region pools)
Fault-tolerant, faster startup Unreachable Bedrock regions or Polly engines no longer abort startup; they are skipped with a warning and retried on the next refresh — and startup is faster overall
Usage & cost tracking Request logs report the usage actually billed by AWS with its cost computed from live AWS pricing, optionally exported as CloudWatch metrics (CLOUDWATCH_METRICS); the previous token estimation is removed and its TOKENS_ESTIMATION* settings are deprecated and ignored
Cost attribution Request, server, and user correlation metadata is now attached to every synchronous Bedrock inference call — the InvokeModel family included, not only Converse — so Bedrock invocation logs can be filtered and costs attributed per request or per user
Smaller container images The published images shrink by around 40% — 413 MB to 253 MB for the AWS Marketplace image, 377 MB to 230 MB for the community image — cutting pull time and storage. ffmpeg is now built with only the audio encoders the server uses, and the unused OpenTelemetry gRPC exporter is no longer installed
Video retention (AWS_S3_VIDEOS_EXPIRES_AFTER) Optional retention period for generated videos, reported as expires_at and enforced on download
Upload expiry (expires_after) Multipart upload sessions honor the OpenAI expires_after policy on the resulting file
Session storage encryption Optional KMS key for Amazon Bedrock session storage (AWS_BEDROCK_SESSION_ENCRYPTION_KEY_ARN)
Proxy trust (PROXY_TRUSTED_HOSTS) X-Forwarded-* headers are only honored when sent by a trusted reverse-proxy address
Input file size limit (MAX_INPUT_FILE_SIZE) Optional cap on the size of downloaded/decoded input files, with bounded download concurrency (MAX_CONCURRENT_INPUT_DOWNLOADS)
Legacy model opt-in fix AWS_BEDROCK_LEGACY now also exposes models whose AWS legacy date has already passed (e.g. Amazon Nova Reel)

Security Hardening

  • MCP transports and /search_models now require authentication when an API key is configured — clients that relied on these endpoints being open must now send the API key
  • SSRF protection hardened against IP-literal encoding and DNS-rebinding bypasses on URL file inputs
  • s3:// file inputs are restricted to the server's allowed buckets, and multipart upload filenames are validated
  • Decoded image size is capped against decompression-bomb payloads
  • ARNs and AWS account IDs are redacted from client-facing error messages, and presigned URL signatures are stripped from logs and traces
  • An empty resolved API key (e.g. a blank secret value) now disables authentication cleanly instead of matching an empty bearer token, and CORS no longer allows credentialed cross-origin requests
  • Reduced container attack surface: the images no longer carry the video codec, X11 and font libraries that a distribution ffmpeg package links — x264, x265, AOM, dav1d, SVT-AV1 and SDL2 among them — nor the gRPC stack. None were reachable from the audio transcoding ffmpeg is used for, and they accounted for the bulk of the images' third-party native code. ffmpeg is built from the same version the base distribution ships, with only the audio encoders in use, and enables no GPL-licensed component

Agent SDK Compatibility

The Responses API request/response surface was audited and hardened against the OpenAI SDK and real agent clients, end-to-end tested against the OpenAI Codex CLI:

  • Hosted and agent tool types (web_search, computer_use, file_search, custom/namespace tools, and other items without a Bedrock equivalent) are now accepted and dropped instead of rejected with 400, preserving compatibility with existing agent tooling
  • Streaming responses now correctly terminate with response.incomplete or response.failed (matching upstream behavior) instead of always reporting response.completed
  • Mid-stream errors emit the spec-compliant error SSE event
  • input_tokens usage now includes cache read/write tokens, matching OpenAI's accounting
  • url_citation annotations are emitted alongside their streaming events
  • Echoed reasoning items tolerate the field variations produced by different SDKs and agent clients

Fixes

  • Rerank models are no longer incorrectly advertised on Converse-based chat routes and MCP tools
  • Fixed per-request model parameter overrides (default_model_params) occasionally leaking into subsequent requests for the same model
  • 5xx provider errors now report server-side error types (server_error/api_error) in OpenAI and Anthropic error envelopes instead of invalid_request_error
  • Unknown paths (404) and wrong methods (405) now return the error envelope of the API family they were sent to, instead of the framework's default detail payload
  • The Anthropic Messages API now returns 404 instead of 400 for an unknown model, matching the upstream API, and rejects a top_p above 1.0
  • Audio transcription returns plain text for response_format=text and now defaults verbose_json to segment timestamps
  • Responses API usage reports input_tokens_details.cache_write_tokens, which recent OpenAI SDKs require to parse a response
  • Files API listing and cursor pagination order by creation time again: file IDs now use an order-preserving alphabet, where the previous one could sort a newer file first. IDs issued before this release keep working, but sort among themselves as before until they expire
  • Newer Anthropic client request fields (free-form JSON Schema keywords in tool input_schema, adaptive thinking display) are accepted instead of rejected in strict validation mode
  • Amazon Nova 2 no longer fails on max_tokens combined with high reasoning effort (the cap is dropped with a logged warning)
  • Explicit cache points are kept off tool-related content blocks for models without tool caching support
  • The Files API unavailable error no longer exposes the S3 bucket configuration detail
  • Fixed input files from one request occasionally leaking into later requests served by the same connection, which could fail those requests with internal errors
  • Anthropic Messages streams now emit an empty tool-input delta for tool calls without arguments, so SDK stream accumulators no longer fail on argument-less tool calls
  • JSON-body image edit and variation requests now accept the model field instead of rejecting the request
  • Model listings now report service: "AWS Bedrock Runtime" for classic Bedrock models (previously "Amazon Bedrock"), distinguishing them from "AWS Bedrock Mantle"
  • High reasoning effort now maps to the intended thinking-token budget on Anthropic Claude models (the budget factor was previously miscomputed)
  • Setting log_level to disabled now suppresses all log output as documented, instead of publishing every event
  • Server startup no longer fails when the ECS container metadata endpoint answers slowly, which could prevent small Fargate tasks from starting: the lookup is retried, then falls back to the STS caller identity with a startup warning
  • Multipart upload parts are numbered from the parts already stored in S3 instead of a per-instance counter: with several server instances behind a load balancer, two parts of one upload could be given the same number, overwriting each other and failing the upload
  • Multi-region failover now covers a region that does not offer the service at all: with no AWS_COMPREHEND_REGION set, a Bedrock region without Amazon Comprehend moves on to the next one as documented, instead of failing language detection and Comprehend moderation

v1.13.0 – 2026-07-03 – Terraform Module Compliance & Security Hardening

At a glance

  • A Terraform-module release — the stdapi-ai module and its VPC, KMS and ECS Fargate children, with no server change.
  • Security Hub FSBP control documentation — every module README now carries a full Foundational Security Best Practices control mapping.
  • Compliance gaps closed — default security group lockdown, ALB access logging, and EFS POSIX user enforcement with native backups.
  • Optional network integrations — compliance VPC endpoints, a GuardDuty VPC endpoint, Route 53 Resolver DNS Firewall, and VPC Flow Logs retention.
  • Tagging and token cost — all four modules accept a tags variable, and MCP tool descriptions were shrunk, lowering the token cost of every agent session.

This release focuses on the stdapi-ai Terraform module and its child modules — VPC, KMS, and ECS Fargate — adding detailed AWS Security Hub control documentation and closing several compliance gaps: default security group lockdown, ALB access logging, EFS POSIX user enforcement with native backups, and optional compliance/GuardDuty/DNS Firewall VPC integrations. All four modules now also accept a tags variable for custom resource tagging.

Documentation-first release

Every module README now includes a full Security Hub Foundational Security Best Practices (FSBP) control mapping. See Authentication & Security for a summary and links to each module.

Fixes

  • Added the missing 1h and 5m values to PromptCacheRetention for Bedrock-specific prompt cache TTLs in the OpenAI Responses API

Security Hub & Compliance Hardening

Feature Module Description
Security Hub FSBP control documentation VPC, KMS, ECS Fargate, stdapi-ai Per-control (pass/fail/conditional/N-A) tables added to each module README
Default security group lockdown VPC New aws_default_security_group resource revokes all default ingress/egress rules (EC2.2 / CIS 5.4)
VPC Flow Logs retention VPC Default retention increased from 7 to 365 days (EC2.6)
Compliance VPC endpoints VPC New compliance_vpc_endpoints_enabled variable adds ECR, SSM, SSM Contacts, and SSM Incidents interface endpoints
GuardDuty VPC endpoint VPC New guardduty_vpc_endpoint_enabled variable adds the guardduty-data interface endpoint
Route 53 Resolver DNS Firewall VPC New dns_firewall_enabled variable blocks/alerts on DNS queries to known-malicious domains (AWS Managed Domain Lists, plus DGA/DNS-tunneling detection via dns_firewall_advanced_enabled); dedicated VPC only
ALB access logging stdapi-ai New alb_access_logging_enabled variable (default true) logs ALB access to a dedicated, encrypted S3 bucket
EFS POSIX user enforcement ECS Fargate mount_points now accepts an efs_posix_user object to enforce a POSIX identity on EFS access points (EFS.4)
EFS native backups ECS Fargate New mount_points_efs_backup_enable variable enables native EFS automatic backups, independent of the existing AWS Backup plan (EFS.7)
Resource tagging VPC, KMS, ECS Fargate, stdapi-ai New tags variable propagates custom tags to nearly all created resources (IAM.24 / EC2.48)

Other Infrastructure Changes

Feature Description
AWS provider version bump Requirement raised to >= 6.27.0 across all four modules
S3 object tag rename Files API objects and the corresponding Terraform lifecycle rule now use the stdapi-ai.expires tag key instead of expires; a temporary backward-compatible rule still expires legacy-tagged objects
aws-apn-id resource tagging AWS resources created at runtime (Bedrock async jobs, Transcribe jobs, S3 objects) are tagged with aws-apn-id, the standard AWS Marketplace attribution tag — an internal, vendor-side tag, not user-configurable

MCP Token Optimization

  • Significantly reduced the size of MCP tool descriptions across the API, lowering the token cost of every AI agent session connected to this server
  • No change in functionality: all parameter constraints and usage guidance remain intact

v1.12.0 – 2026-05-29 – Completions API, Video Understanding & File References

At a glance

  • /v1/completions — the OpenAI text completion endpoint, for text-first coding agents and legacy completion clients.
  • Video understanding — TwelveLabs Pegasus analyses video/* inputs in chat completions, honouring service_tier and guardrail configuration.
  • Input token counting — /v1/responses/input_tokens for the Responses API.
  • The file-id: URI scheme — reference a Files API upload anywhere a URL is accepted: embeddings, transcription, chat, images and messages.
  • Settings and compatibility — DEFAULT_MODEL_SERVICE_TIERS applies a per-model service tier automatically, reasoning can be explicitly enabled or disabled, and the Anthropic /v1/messages route accepts system-role messages.

This release adds the OpenAI-compatible /v1/completions endpoint for text-first coding agents and legacy completion clients, TwelveLabs Pegasus video understanding for analyzing video/* inputs in chat completions, and an input token counting endpoint for the Responses API. Files uploaded through the Files API can now be referenced anywhere a URL is accepted using the new file-id: URI scheme. The Anthropic Messages API now accepts system-role messages (merged into the system prompt for compatibility), reasoning can be explicitly enabled or disabled, and a new DEFAULT_MODEL_SERVICE_TIERS setting applies per-model service tiers automatically.

Chat Completions

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/completions – text completion endpoint for text-first coding agents Amazon Bedrock Amazon Bedrock - foundation models
OpenAI OpenAI /v1/responses/input_tokens – input token counting Amazon Bedrock Amazon Bedrock - CountTokens API
Anthropic Anthropic /v1/messages – accepts system-role messages (merged into the system prompt) Amazon Bedrock Amazon Bedrock - Claude models
Twelve Labs Twelve Labs Pegasus video understanding (video/* inputs) Amazon Bedrock Amazon Bedrock - TwelveLabs Pegasus

Speech & Audio

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/audio/speech – case-insensitive voice names & default model Amazon Polly Amazon Polly

Platform Features

Feature Description
file-id: URI scheme Reference Files API uploads via file-id:<file-id> anywhere a URL is accepted — embeddings, audio transcription/translation, chat, images, and messages
Default model service tiers (DEFAULT_MODEL_SERVICE_TIERS) Automatically apply a per-model service tier (default, flex, priority, reserved) when none is provided in the request
Explicit reasoning enable/disable Reasoning/thinking can now be explicitly enabled or disabled via request parameters
Service tier & guardrail support for Pegasus TwelveLabs Pegasus requests honor service_tier and Bedrock Guardrail configuration
MCP speech streaming defaults to SSE /v1/audio/speech defaults stream_format to sse when invoked as an MCP tool for broader client compatibility
Full regional S3 bucket handling The Terraform module resolves regional S3 buckets via resource-level region (requires AWS provider >= 6.0.0)
Reliable cross-region model identifiers Region routing no longer fails intermittently with "The provided model identifier is invalid": a region whose inference profile is missing or not yet propagated is skipped, and a geo-scoped profile is never sent to a different region

v1.11.0 – 2026-05-02 – MCP Server, Agent Discovery & Model Search (with v1.11.1–v1.11.4 maintenance updates, through 2026-05-28)

At a glance

  • An MCP server — every API endpoint is exposed as an MCP tool, over Streamable HTTP and SSE transports that are independently enabled and selectively restricted.
  • /search_models — filter models by route, MCP tool name, input and output modalities, region, streaming support and legacy status.
  • Agent discovery — RFC 8288 Link headers on /, an RFC 9727 API catalog at /.well-known/api-catalog, an MCP Server Card, and robots.txt content signals.
  • JSON bodies for binary endpoints — audio transcription, audio translation and image edits accept application/json with files as base64, data URI, HTTP URL or S3 URI.
  • Maintenance (v1.11.1–v1.11.4) — max_tokens made optional on Anthropic /v1/messages, MCP dependencies added to the container image, and Starlette upgraded for CVE-2026-48710.

This release introduces a Model Context Protocol (MCP) server, making all stdapi.ai API endpoints directly accessible as MCP tools for AI agents and agentic workflows. A new /search_models endpoint enables precise discovery of models by route, MCP tool, region, streaming support, and legacy status. Agent-friendly discovery metadata is now exposed via RFC 8288 Link headers and an RFC 9727 machine-readable API catalog at /.well-known/api-catalog. Endpoints that previously required binary multipart/form-data uploads now also accept an application/json body for MCP and HTTP client compatibility. The Anthropic Messages API now accepts xhigh as a reasoning_effort value.

MCP Server

Feature Description
MCP server (Streamable HTTP & SSE) All API endpoints exposed as MCP tools; Streamable HTTP and SSE transports can be independently enabled or disabled via configuration
Configurable MCP tool exposure Individual MCP tools can be selectively enabled or restricted via configuration
JSON body for binary endpoints Audio transcription, audio translation, and image edit endpoints now accept application/json with files as base64, data URI, HTTP URL, or S3 URI
Feature Description
/search_models New official endpoint to filter models by route, MCP tool name, input/output modalities, region, streaming, and legacy status; returns richer metadata than /v1/models or Anthropic /v1/models, designed for LLM-driven model selection (replaces BETA and undocumented /available_models)

Agent Discovery

Feature Description
RFC 8288 Link headers Root (/) endpoint returns Link headers for resource discovery
RFC 9727 API catalog (/.well-known/api-catalog) Machine-readable API catalog for automated agent and tool discovery
MCP Server Card (/.well-known/mcp/server-card.json) Advertises available MCP transports and capabilities to AI agents (SEP-1649)
robots.txt AI signals Updated robots.txt with Content-Signal directives and explicit /.well-known/ allow rule

Chat Completions & Messages

Provider Endpoint/Feature AWS Backend
Anthropic Anthropic /v1/messages reasoning_effort=xhigh support Amazon Bedrock Amazon Bedrock - Claude models

Deprecation Mappings

  • Added automatic fallback for amazon.nova-reel-v1:0 and anthropic.claude-3-haiku-20240307-v1:0 to their respective replacements

Fixes

  • Fix reasoning token double-counting in usage calculation in OpenAI Responses API adapter
  • Fix missing file_id inputs for image and file processing in OpenAI Responses API adapter
  • Remove store parameter from unsupported validations in chat completions to ensure client compatibility

Fixes & Maintenance (v1.11.1–v1.11.4)

v1.11.1

  • Make max_tokens optional in Anthropic /v1/messages to align with the Anthropic API specification
  • Remove unsupported reasoning configuration checks for broader client compatibility
  • Rename /v1/responses route tag from "Responses" to "Chat" in OpenAPI documentation for consistency

v1.11.2-v1.11.3

  • Add missing MCP dependencies to container image.

v1.11.4

  • Upgrade Starlette dependency to fix CVE-2026-48710.

v1.10.0 – 2026-04-17 – OpenAI Responses API

At a glance

  • /v1/responses — OpenAI's API for agents and multi-step workflows, drop-in compatible with the OpenAI SDK.
  • Every Converse-compatible model — it works with all Amazon Bedrock Converse-compatible models, streaming included.
  • Built-in tools — web_search / web_search_preview, code_interpreter and image_generation.
  • Function tools, extended reasoning and structured output on the same surface.
  • Fixes — prompt caching with tool-related content, an optional signature field in Anthropic message types, and model legacy detection when the end-of-life date falls before the next cache refresh.

This release adds support for the OpenAI /v1/responses endpoint—OpenAI's next-generation API designed for building agents and multi-step AI workflows. Drop-in compatible with the OpenAI SDK, it works with all Amazon Bedrock Converse-compatible models and supports streaming, function tools, built-in tools (web search, code interpreter, image generation), extended reasoning, and structured output.

Responses (OpenAI-Compatible)

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/responses Amazon Bedrock Amazon Bedrock - foundation models
OpenAI OpenAI /v1/responses – web_search / web_search_preview built-in tool Amazon Nova Amazon Nova models
OpenAI OpenAI /v1/responses – code_interpreter built-in tool Amazon Nova Amazon Nova models
OpenAI OpenAI /v1/responses – image_generation built-in tool Amazon Bedrock Amazon Bedrock - image models

Fixes

  • Fix prompt caching error when messages contain tool-related content on models that do not support tool caching
  • Make signature field optional in Anthropic message types
  • Fix model legacy detection when the end-of-life date falls before the next cache refresh

v1.9.0 – 2026-04-10 – Files API & Images API JSON Body

At a glance

  • A Files API backed by Amazon S3 — /v1/files CRUD on both the OpenAI-compatible and Anthropic-compatible interfaces, sharing one store.
  • Incremental uploads — /v1/uploads, the OpenAI multipart uploads API, for large files.
  • File IDs as model inputs — a stored file is usable as a document or image input in chat completions and in messages.
  • A JSON body for image editing — /v1/images/edits and /v1/images/variations accept application/json referencing Files API IDs or URLs, so pipeline steps chain without re-uploading.
  • New required configuration — AWS_S3_BUCKET must be set, with read, write, delete and list permissions on it.

This release introduces a Files API backed by Amazon S3, available through both the OpenAI-compatible and Anthropic-compatible interfaces. Files uploaded via either API share the same S3 storage and can be referenced across both interfaces. Large files can be uploaded incrementally using the OpenAI multipart uploads API. Stored files can be referenced by ID directly in image edit and variation requests (JSON body), as well as in chat completion messages as document or image inputs. The image edits endpoint now also accepts an application/json body as an alternative to multipart form-data, making it easier to chain pipeline steps without re-uploading files.

New Required Configuration

Files API requires AWS_S3_BUCKET to be configured (shared with the image URL response feature). The S3 prefix for stored files defaults to files/ and is configurable via AWS_S3_FILES_PREFIX. Ensure your IAM role includes read, write, delete, and list permissions on the files prefix in addition to the existing S3 permissions for presigned URLs.

Files & Storage

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/files – CRUD operations Amazon S3 Amazon S3
OpenAI OpenAI /v1/uploads – multipart uploads Amazon S3 Amazon S3
Anthropic Anthropic /v1/files – CRUD operations Amazon S3 Amazon S3

Image Generation

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/images/edits – JSON body with images/mask referencing Files API IDs or URLs Amazon Bedrock Amazon Bedrock - image models
OpenAI OpenAI /v1/images/variations – JSON body with image referencing a Files API ID or URL Amazon Bedrock Amazon Bedrock - image models

Chat Completions & Messages

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI Files API file IDs usable as document/image inputs in chat completions Amazon Bedrock Amazon Bedrock - foundation models
Anthropic Anthropic Files API file IDs usable as document/image inputs in messages Amazon Bedrock Amazon Bedrock - foundation models

Fixes

  • Document inputs via S3 URLs are not supported as Bedrock Converse API inputs for some models (e.g., Claude) — now properly detected and handled

v1.8.0 – 2026-04-04 – Broader Model Compatibility & Structured Output

At a glance

  • Structured output — response_format with JSON object and JSON schema on OpenAI chat completions.
  • Request metadata — metadata is forwarded to Bedrock, and the request context (request_id, server_id, user_id) is tagged onto every Bedrock and Amazon Transcribe job.
  • Tool handling — Amazon Nova's grounding tool maps to web_search content blocks, with multi-turn support, and the broken systemTool_ auto-promotion was removed.
  • Region routing — region-restricted models always get non-global inference profiles, and the case where no region is usable is handled gracefully.
  • New required IAM permissions — bedrock:TagResource and transcribe:TagResource; AWS_BEDROCK_LEGACY now defaults to false.

This release focuses on improving reliability and compatibility across a wide variety of models. Structured response formats (JSON object and JSON schema) are now supported on OpenAI chat completions, and request metadata can be forwarded to Bedrock. Tool handling has been significantly improved—both for model-specific system tools and for Amazon Nova's grounding tool, including multi-turn support. Region routing is now more robust, correctly enforcing non-global inference profiles for region-restricted models and handling edge cases gracefully.

New Required IAM Permissions

v1.8.0 requires two new IAM permissions to attach request metadata tags to jobs:

  • bedrock:TagResource on arn:aws:bedrock:*:*:async-invoke/* — needed for Bedrock asynchronous invocation jobs (see IAM Permissions). The twelvelabs.marengo-embed-3-0-v1:0 and twelvelabs.marengo-embed-2-7-v1:0 models rely on asynchronous invocation and will fail with an access denied error if this permission is missing.
  • transcribe:TagResource on arn:aws:transcribe:*:*:transcription-job/* — needed for Amazon Transcribe transcription jobs (see IAM Permissions). The amazon.transcribe model will fail with an access denied error if this permission is missing.

Ensure your IAM role or user policy includes both statements before upgrading to v1.8.0.

Chat Completions

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI response_format – JSON object and JSON schema structured output Amazon Bedrock Amazon Bedrock - foundation models
OpenAI OpenAI metadata – request metadata forwarding to Bedrock Amazon Bedrock Amazon Bedrock - foundation models
Amazon Nova Amazon Nova Nova Code Interpreter global profile support Amazon Bedrock Amazon Bedrock - Nova models

Messages (Anthropic-Compatible)

Provider Endpoint/Feature AWS Backend
Amazon Nova Amazon Nova nova_grounding responses mapped to web_search content blocks Amazon Bedrock Amazon Bedrock - Nova models
Amazon Nova Amazon Nova Multi-turn conversation support with nova_grounding Amazon Bedrock Amazon Bedrock - Nova models

Platform Features

Feature Description
Non-global profiles for region-restricted models Region-restricted models are now always assigned non-global inference profiles, preventing requests from bypassing configured region restrictions
Region routing edge case handling Region routing gracefully handles cases where no usable regions are available
ECS-based server ID When running on ECS, server_id in logs is set to task_id.container_name for precise instance identification across tasks and containers
Request metadata tagging stdapi.ai request context (request_id, server_id, user_id) is automatically attached as tags to every Bedrock and Amazon Transcribe job, making it easy to trace API calls across AWS service logs

Fixes

  • Fix systemTool_ prefix handling: removed broken auto-promotion logic; system tools require specific tool output handling not compatible with generic tool forwarding
  • AWS_BEDROCK_LEGACY default changed from true to false to prevent access denied errors on legacy models that have not been actively used recently
  • Bedrock read timeouts are now handled as standard model errors (503) instead of unhandled exceptions, and are properly retried across regions when multi-region routing is enabled

v1.7.0 – 2026-03-20 – Automatic Region Routing, Deprecated Model Fallback & Resilience Improvements

At a glance

  • Automatic multi-region routing — Bedrock requests are distributed across the configured AWS regions, failing over on quota limits or unavailability.
  • More quota by adding regions — each region carries its own quota, so adding one multiplies the effective tokens-per-minute and daily limits.
  • Deprecated model fallback — deprecated model IDs are transparently rerouted to their replacements, with an extensible mapping, so clients survive AWS model retirements unchanged.
  • A configurable AI response timeout, so a model call cannot hang indefinitely.
  • S3 URLs for file inputs across all relevant endpoints, alongside HTTP URLs, data URIs and base64, with memory-efficiency improvements.

The headline feature of v1.7 is automatic multi-region routing: stdapi.ai now intelligently distributes requests across your configured AWS regions, failing over automatically on quota limits or unavailability—and because each region carries its own independent quota, adding regions directly multiplies your effective tokens-per-minute and daily limits. Alongside this, deprecated model IDs are transparently redirected to their replacements so clients survive AWS model retirements without any code changes. This release also adds S3 URL support for file inputs across all relevant endpoints, a configurable AI response timeout, and memory efficiency improvements.

Platform Features

Feature Description
Automatic region routing with configurable strategies Intelligently distributes Bedrock requests across configured AWS regions with automatic failover on quota limits or unavailability; supports ordered, lowest_latency, and round_robin strategies
Deprecated model fallback Transparently reroute deprecated model IDs to their replacements; extend or override the built-in mapping; warns on legacy model usage
AI response timeout Configurable timeout for AI model responses to prevent indefinitely hanging requests
Expanded file input support File inputs (images, documents, audio) now support S3 URLs in addition to HTTP URLs, data URIs, and plain base64 across all relevant endpoints; improves memory efficiency by releasing file data as early as possible
Model lifecycle timestamps Model created/updated timestamps now derived from lifecycle data (startOfLifeTime, endOfLifeTime)

Fixes

  • Fix SSE stream error handling in monitoring to handle specific API and AWS client errors gracefully
  • Fix audio MIME type detection failure when libmagic's in-memory buffer path silently returns application/octet-stream; fall back to file-based detection to ensure correct format is sent to Bedrock

v1.6.0 – 2026-02-27 – Anthropic API Compatibility & Advanced Claude Capabilities

At a glance

  • A full Anthropic-compatible API — /v1/messages and /v1/messages/count_tokens, usable straight from the Anthropic SDK.
  • Anthropic-format model discovery — /v1/models and /v1/models/{model_id}.
  • Claude server tools — bash, text editor, computer and memory, on both the Messages surface and OpenAI chat completions; Amazon Nova's web_search maps to nova_grounding.
  • Configurable route prefixes — ANTHROPIC_ROUTES_PREFIX and OPENAI_ROUTES_PREFIX, plus Anthropic beta flag filtering to prevent Bedrock ValidationException errors.
  • Real usage tracking — token counts sourced directly from AWS billing data instead of tiktoken estimation, and Claude model name aliases resolved to Bedrock identifiers.

Introduces a full Anthropic-compatible API layer, enabling direct use of the Anthropic SDK and Claude-native tools with Amazon Bedrock. Adds Claude server tools support via OpenAI chat completions, token count estimation, automatic Anthropic beta flag filtering, and configurable route prefixes.

Chat Completions

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/chat/completions Claude server tools (bash, str_replace_based_edit_tool, computer, memory) Claude Claude models on Amazon Bedrock

Messages (Anthropic-Compatible)

Provider Endpoint/Feature AWS Backend
Anthropic Anthropic /v1/messages – Full Anthropic Messages API Amazon Bedrock Amazon Bedrock - Converse API
Anthropic Anthropic /v1/messages/count_tokens – Token counting Amazon Bedrock Amazon Bedrock - CountTokens API
Claude Claude Claude server tools (bash, text editor, computer, memory) Amazon Bedrock Amazon Bedrock - Claude models
Amazon Nova Amazon Nova Web search tool (web_search → nova_grounding) Amazon Bedrock Amazon Bedrock - Nova models

Model Discovery (Anthropic-Compatible)

Provider Endpoint/Feature AWS Backend
Anthropic Anthropic /v1/models – List models (Anthropic format) Amazon Bedrock Amazon Bedrock - model catalog
Anthropic Anthropic /v1/models/{model_id} – Get model details Amazon Bedrock Amazon Bedrock - model catalog

Platform Features

Feature Description
ANTHROPIC_ROUTES_PREFIX configuration Configurable base path prefix for Anthropic-compatible routes (default: /anthropic)
OPENAI_ROUTES_PREFIX configuration Configurable base path prefix for OpenAI-compatible routes
Real usage tracking (usage in logs) Token counts sourced directly from AWS billing data (replaces tiktoken-based estimation)
Anthropic beta flag filtering (ANTHROPIC_BETA_FILTER) Automatically filter unsupported anthropic-beta flags to prevent Bedrock ValidationException errors; extensible via ANTHROPIC_BETA_ALLOWLIST
Claude model name aliases Use official Anthropic model names (e.g., claude-opus-4-8) auto-resolved to Amazon Bedrock identifiers

v1.5.0 – 2026-02-15 – Advanced Reasoning & Model Compatibility (with v1.5.1–v1.5.2 maintenance updates, through 2026-02-18)

At a glance

  • Amazon Nova 2 reasoning implemented on chat completions.
  • Claude 4.6+ adaptive reasoning configuration.
  • System prompt handling for models that do not support one, widening model compatibility.
  • v1.5.1 — Amazon Nova Canvas image editing falls back to the TEXT_IMAGE task type when no mask is provided.
  • v1.5.2 — a / route so the root endpoint stops answering 404, and empty system content blocks are handled for Converse API compatibility.

Introduces advanced reasoning capabilities with Amazon Nova 2 and Anthropic Claude 4.6+ adaptive reasoning, enhanced system prompt handling for broader model compatibility.

Chat Completions

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI System prompt handling for unsupported models Amazon Bedrock Amazon Bedrock - foundation models
Amazon Nova Amazon Nova Nova 2 chat model reasoning implementation Amazon Bedrock Amazon Bedrock - foundation models
Claude Claude Claude 4.6+ adaptive reasoning configuration Amazon Bedrock Amazon Bedrock - Claude models

Fixes & Maintenance (v1.5.1–v1.5.2)

v1.5.2

  • Add "/" route to avoid 404 errors on root endpoint
  • Fix empty system content block handling (improves Amazon Bedrock Converse API compatibility)

v1.5.1

  • Fix Amazon Nova Canvas image editing to fall back to TEXT_IMAGE task type when no mask is provided

v1.4.0 – 2026-02-11 – Audio Enhancements & Model Compatibility

At a glance

  • Mistral Voxtral joins the audio models.
  • Speaker diarization — the diarized_json format on /v1/audio/transcriptions.
  • Audio formats on chat completions, with extended Bedrock finish-reason mapping.
  • Prompt caching TTL support on chat completions.
  • Model aliasing — OpenAI-style model names resolved for seamless compatibility.

Expands audio capabilities with Mistral Voxtral support, speaker diarization, audio formats for chat completions, and introduces prompt caching TTL and model aliasing for better OpenAI compatibility.

Chat Completions

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/chat/completions audio format support Amazon Bedrock Amazon Bedrock - foundation models
OpenAI OpenAI /v1/chat/completions extended Bedrock finish reasons mapping Amazon Bedrock Amazon Bedrock
OpenAI OpenAI Prompt caching TTL support Amazon Bedrock Amazon Bedrock - prompt caching

Speech & Audio

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/audio/transcriptions diarized_json format Amazon Transcribe Amazon Transcribe
Mistral Mistral Voxtral audio model Amazon Bedrock Amazon Bedrock - foundation models

Platform Features

Feature Description
Model alias support Seamless OpenAI compatibility via model name aliasing

Fixes

  • Fix chat completion file input handling and refactor base64 decoding and MIME handling for file processing.
  • Re-raise startup exceptions and disable botocore logging to improve error visibility

v1.3.0 – 2026-01-11 – Image Editing & Variation Support (with v1.3.1–v1.3.5 maintenance updates, through 2026-02-02)

At a glance

  • /v1/images/edits — OpenAI-compatible image editing backed by Amazon Bedrock.
  • /v1/images/variations — OpenAI-compatible image variations.
  • DEFAULT_TTS_LANGUAGE (v1.3.2) — a configurable default language for text-to-speech, plus image[] array-style notation for image edits.
  • Tool-call and streaming fixes (v1.3.1, v1.3.3–v1.3.5) — robust JSON parsing and validation of tool arguments, no premature contentBlockStop in streamed chat completions, and empty content blocks skipped in assistant responses.
  • A deprecation mapping (v1.3.4) — amazon.titan-image-generator-v2:0 to amazon.nova-canvas-v1:0.

Adds support for OpenAI's image editing and variation endpoints, enabling image manipulation capabilities backed by Amazon Bedrock. Includes maintenance updates for content block handling, tool call validation, streaming fixes, and TTS optimization.

Image Generation

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/images/edits Amazon Bedrock Amazon Bedrock - image models
OpenAI OpenAI /v1/images/variations Amazon Bedrock Amazon Bedrock - image models

Speech & Audio (v1.3.2)

Feature Description
DEFAULT_TTS_LANGUAGE setting Configurable default language for TTS to optimize performance

Fixes & Maintenance (v1.3.1–v1.3.5)

v1.3.5

  • Refactor content block handling to skip empty entries in assistant responses

v1.3.4

  • Handle invalid tool call arguments with robust JSON content validation
  • Add deprecation mapping for amazon.titan-image-generator-v2:0 → amazon.nova-canvas-v1:0

v1.3.3

  • Remove premature stop condition for contentBlockStop in streaming chat completions

v1.3.2

  • Support image[] array-style notation for OpenAI image edits
  • Handle empty audio segments in transcription duration calculation

v1.3.1

  • Improve JSON parsing for tool arguments and results
  • Correct example → examples in OpenAPI model path parameter

v1.2.0 – 2025-12-18 – Service Tiers, System Tools & Performance Enhancements

At a glance

  • Service tiers — the service_tier parameter on chat completions, with latency headers on all Bedrock routes.
  • Bedrock-specific system tools — Amazon Nova grounding on /v1/chat/completions.
  • GPT5.2 API update — reasoning_effort=xhigh accepted.
  • A configuration flag for guardrail override allow, controlling what a request may override on Amazon Bedrock Guardrails.
  • Python 3.14 with performance optimization, and direct aiobotocore usage replacing aioboto3.

Introduces service tiers and latency headers for all Bedrock routes, Bedrock-specific system tools (Nova grounding), GPT5.2 API compatibility, configurable guardrail overrides, and Python 3.14 optimization.

Chat Completions

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/chat/completions service_tier parameter Amazon Bedrock Amazon Bedrock - service tiers
OpenAI OpenAI /v1/chat/completions Bedrock-specific system tools (Nova grounding) Amazon Bedrock Amazon Bedrock - system tools
OpenAI OpenAI /v1/chat/completions GPT5.2 API update (reasoning_effort=xhigh)

Content Safety & Moderation

Feature AWS Backend
Configuration flag for guardrail override allow Amazon Bedrock Amazon Bedrock Guardrails

Platform Features

Feature AWS Backend / Description
Service tiers and latency headers (all Bedrock routes) Amazon Bedrock Amazon Bedrock - service tiers
Python 3.14 support Upgraded to Python 3.14 with performance optimization
Dependency update Direct aiobotocore usage (replaced aioboto3)

Fixes

  • Fix warnings for duplicated FastAPI routes (/docs and /openapi.json).

v1.1.0 – 2025-11-27 – Embeddings Enhancement, Prompt Caching & Advanced Routing

At a glance

  • Multimodal embeddings — Amazon Nova multimodal models and TwelveLabs Marengo V3.
  • Intelligent embedding plumbing — S3 multimodal upload and sync/async Bedrock invocation chosen per request.
  • Prompt caching — prompt_cache_key on /v1/chat/completions, plus the GPT5.1 API update (reasoning_effort=none).
  • Advanced routing — Amazon Bedrock application inference profiles and prompt routers.
  • ARN handling — server-side ARN mapping, with optional client-side ARN passing.

Expands multimodal embedding capabilities, adds prompt caching support, and introduces advanced routing with application inference profiles and prompt routers.

Chat Completions

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI Prompt caching /v1/chat/completions prompt_cache_key Amazon Bedrock Amazon Bedrock - prompt caching
OpenAI OpenAI /v1/chat/completions GPT5.1 API update (reasoning_effort=none)

Embeddings

Provider Endpoint/Feature AWS Backend
Intelligent S3 multimodal upload Amazon S3 Amazon S3
Intelligent Sync/async Bedrock invocation Amazon Bedrock Amazon Bedrock
Amazon Nova Amazon Nova Multimodal embeddings models
Twelve Labs Twelve Labs Marengo V3 models

Advanced Routing

Feature AWS Backend
Application inference profiles Amazon Bedrock Amazon Bedrock - application inference profiles
Prompt routers Amazon Bedrock Amazon Bedrock - prompt routers
Server-side ARN mapping Amazon Bedrock Amazon Bedrock
Client-side ARN passing (optional) Amazon Bedrock Amazon Bedrock

Fixes

  • /v1/chat/completions: Fix default value passed to the converse API for tools without parameters.
  • stdapi-ai Terraform module: Fix error if alarms_enabled = true but sns_topic_arn undefined.

v1.0.0 – 2025-11-10 – Foundation Release

At a glance

  • Chat completions — /v1/chat/completions on every model supporting the Converse and ConverseStream APIs, with DeepSeek reasoning_content and Qwen thinking parameters.
  • Embeddings — /v1/embeddings on Cohere Embed V3 and V4, TwelveLabs Marengo V2, and Amazon Titan Embed V1 and V2.
  • Speech and audio — /v1/audio/speech, /v1/audio/transcriptions and /v1/audio/translations, on Amazon Polly, Amazon Transcribe and Amazon Translate.
  • Image generation — /v1/images/generations on Amazon Nova Canvas, Amazon Titan Image Generator and Stability AI models.
  • Platform — Amazon Bedrock Guardrails, cross-region inference and multi-region failover, Amazon S3 file storage, static token authentication from SSM Parameter Store or Secrets Manager, AWS X-Ray tracing and CloudWatch structured logging.

The initial release establishes core OpenAI API compatibility with Amazon Bedrock backing.

Chat Completions

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/chat/completions Amazon Bedrock Amazon Bedrock - foundation models
All models supporting Converse/ConverseStream APIs Amazon Bedrock Amazon Bedrock - Converse API
DeepSeek DeepSeek /v1/chat/completions reasoning_content Amazon Bedrock Amazon Bedrock - foundation models
Qwen Qwen enable_thinking + thinking_budget parameter Amazon Bedrock Amazon Bedrock - foundation models
Qwen Qwen top_k parameter Amazon Bedrock Amazon Bedrock - foundation models

Embeddings

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/embeddings Amazon Bedrock Amazon Bedrock - embedding models
Cohere Cohere Embed V3 & V4 models
Twelve Labs Twelve Labs Marengo V2 models
Amazon Amazon Titan Embed V1 & V2 models

Speech & Audio

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/audio/speech Amazon Polly Amazon Polly + Amazon Comprehend
OpenAI OpenAI /v1/audio/transcriptions Amazon Transcribe Amazon Transcribe
OpenAI OpenAI /v1/audio/translations Amazon Transcribe Amazon Transcribe + Amazon Translate

Image Generation

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/images/generations Amazon Bedrock Amazon Bedrock - image models
Amazon Nova Amazon Nova Canvas V1 models
Amazon Amazon Titan Image Generator V1 & V2 models
Stability AI Stability AI Image Core, Ultra and SD3.5 Large models

Model Discovery

Provider Endpoint/Feature AWS Backend
OpenAI OpenAI /v1/models Amazon Bedrock Amazon Bedrock - model catalog

Platform Features

Feature AWS Backend
Bedrock Features
Content filtering and safety Amazon Bedrock Amazon Bedrock Guardrails
Cross-region inference Amazon Bedrock Amazon Bedrock - global/regional
Application inference profiles Amazon Bedrock Amazon Bedrock - inference profiles
Model parameters (temperature, top_p, etc.) Amazon Bedrock Amazon Bedrock - native parameters
Multi-region failover Amazon Bedrock Amazon Bedrock - multi-region
Bedrock guardrails Amazon Bedrock Amazon Bedrock Guardrails
AWS Services
File storage Amazon S3 Amazon S3 - presigned URLs, Transfer Acceleration
Authentication
Static token authentication AWS Systems Manager AWS SSM Parameter Store / AWS Secrets Manager Secrets Manager
Development mode (no auth)
Observability
Distributed tracing AWS X-Ray AWS X-Ray + OpenTelemetry
Structured logging Amazon CloudWatch Amazon CloudWatch (When running on ECS/EKS)
Health check endpoint
HTTP/Security
CORS support
Trusted host validation
Proxy headers (X-Forwarded-*)
GZip compression
📚 Documentation
Interactive API docs & OpenAPI schema
🔌 Compatibility
Provider-specific parameters