Skip to content

Features — AI Gateway for Amazon Bedrock

stdapi.ai is an AI gateway purpose-built for AWS. It brings OpenAI, Anthropic, Cohere and Ollama API compatibility to Amazon Bedrock and AWS AI services — so the tools, SDKs, and applications your team already uses run against your own AWS account, from the moment they point at a new base URL.


How It Works

stdapi.ai translates OpenAI, Anthropic, Cohere and Ollama API calls into native AWS requests. A tool or SDK that speaks one of the four protocols connects on the base URL alone — no plugins, no custom integrations.

%%{init: {'flowchart': {'htmlLabels': true}} }%%
flowchart LR
  openwebui["<img src='../styles/logo_openwebui.svg' style='height:48px;width:auto;vertical-align:middle;' /> Open WebUI"] --> stdapi["<img src='../styles/logo.svg' style='height:64px;width:auto;vertical-align:middle;' /> stdapi.ai"]
  n8n["<img src='../styles/logo_n8n.svg' style='height:48px;width:auto;vertical-align:middle;' /> n8n"] --> stdapi
  ide["<img src='../styles/logo_vscode.svg' style='height:48px;width:auto;vertical-align:middle;' /> IDE + AI Assistant"] --> stdapi
  openai_app["<img src='../styles/logo_openai.svg' style='height:48px;width:auto;vertical-align:middle;' /> Any OpenAI App"] --> stdapi
  anthropic_app["<img src='../styles/logo_anthropic.svg' style='height:48px;width:auto;vertical-align:middle;' /> Any Anthropic App"] --> stdapi
  stdapi --> bedrock["<img src='../styles/logo_amazon_bedrock.svg' style='height:48px;width:auto;vertical-align:middle;' /> Amazon Bedrock"]
  bedrock --> claude["<img src='../styles/logo_anthropic_claude.svg' style='height:36px;width:auto;vertical-align:middle;' /> Claude"]
  bedrock --> qwen["<img src='../styles/logo_qwen.svg' style='height:36px;width:auto;vertical-align:middle;' /> Qwen"]
  bedrock --> mistral["<img src='../styles/logo_mistralai.svg' style='height:36px;width:auto;vertical-align:middle;' /> Mistral"]
  bedrock --> stability["<img src='../styles/logo_stabilityai.svg' style='height:36px;width:auto;vertical-align:middle;' /> Stability AI"]
  bedrock --> more["✨ and more..."]
  stdapi --> transcribe["<img src='../styles/logo_amazon_transcribe.svg' style='height:48px;width:auto;vertical-align:middle;' /> Amazon Transcribe"]
  stdapi --> polly["<img src='../styles/logo_amazon_polly.svg' style='height:48px;width:auto;vertical-align:middle;' /> Amazon Polly"]
  stdapi --> s3["<img src='../styles/logo_amazon_s3.svg' style='height:48px;width:auto;vertical-align:middle;' /> Amazon S3"]

API Compatibility

100+ endpoints, four protocols, every one on an AWS service

Not a chat proxy with a few extras: chat, retrieval, image, video, audio, batch and file endpoints across all four protocols, each one backed by an AWS service running in your account. Training and account administration stay out of scope — there is no fine-tuning, evaluation or code-interpreter container API, and no Anthropic Admin API — and the Cohere surface covers the two families Amazon Bedrock provides, Embed and Rerank.

What your application calls it for AWS service behind it
Chat completions, Responses, Messages, legacy completions, token counting Amazon Bedrock Converse API · Bedrock Mantle
Server-side conversations and stored responses Amazon Bedrock Sessions · Bedrock Mantle
Embeddings and reranking Amazon Bedrock embedding and rerank models
Vector stores and file search Amazon S3 Vectors · Amazon Bedrock Knowledge Bases
Batch inference Amazon Bedrock batch inference
Image generation, editing and variations Amazon Bedrock image models
Video generation Amazon Bedrock video models
Text-to-speech Amazon Polly
Transcription and speech translation Amazon Transcribe · Amazon Translate · Amazon Nova Sonic
Live speech-to-speech Amazon Bedrock
Content moderation Amazon Bedrock Guardrails · Amazon Comprehend
Files and multipart uploads Amazon S3
Model discovery and pricing Amazon Bedrock · AWS Price List

Anthropic, Cohere and Ollama routes live under /anthropic, /cohere and /ollama, so the four protocols are served side by side without colliding on /v1 — and every prefix can become the path your clients already send (Anthropic, Cohere, Ollama).

Every endpoint, with its parameters

Parameter Coverage

stdapi.ai maps as many parameters as possible to Bedrock equivalents — across all routes, not just chat:

  • Generation controls — temperature, max_tokens, top_p, top_k, stop, seed, frequency_penalty, presence_penalty, logit_bias, top_logprobs, streaming via SSE, token usage reporting
  • Reasoning — reasoning_effort (none/minimal/low/medium/high/xhigh/max), enable_thinking, thinking_budget
  • Tool / function calling — Full OpenAI and Anthropic schemas, parallel tool calls, tool choice modes
  • All content types — System, developer, user, assistant, and tool roles; text, image, audio, video, and document content
  • Response formats — JSON object, JSON schema, streaming chunks, reasoning_content, annotations
  • Model-specific extras — Any parameter beyond the standard API via extra_body or top-level request fields

Bedrock & model differences

Not every parameter maps identically across all models. Check the API documentation for details.


100+ Models Across 20+ Providers

Access every model available on Amazon Bedrock through a single, consistent API — including OpenAI GPT, xAI Grok, and other frontier models. Browse the full list.

  • Claude Anthropic Claude
    Claude Fable/Mythos, Claude Opus, Claude Sonnet, Claude Haiku — including reasoning models. Use official Anthropic model names (e.g., claude-fable-5) — they resolve automatically.

  • OpenAI OpenAI GPT
    GPT frontier models plus open-weight gpt-oss, under OpenAI's own model names. The enrollment-gated Daybreak variants are served and priced with the rest of the family where your account is enrolled with OpenAI's Daybreak programme.

  • Google Google Gemma
    Gemma 4 and other Gemma open-weight variants.

  • Amazon Nova Amazon Nova
    Nova — including reasoning-capable variants. Canvas for images. Multimodal embeddings. Built-in web grounding and code interpreter.

  • Meta Llama Meta Llama
    Llama Scout, Maverick, and earlier Llama variants.

  • Qwen Alibaba Qwen
    Qwen and Qwen3 Coder — including thinking mode.

  • DeepSeek DeepSeek
    Latest DeepSeek V3 models with automatic reasoning content surfacing.

  • Kimi Moonshot Kimi
    Kimi with optional thinking mode.

  • Mistral Mistral AI
    Mistral, Mixtral, and Mistral Large variants.

  • Cohere Cohere
    Embed v3 and v4 for multimodal embeddings; Rerank 3.5 for reranking.

  • Stability AI Stability AI
    Stable Diffusion 3.5, SD3 Ultra, and specialty models (upscale, style, search).

  • MiniMax MiniMax & more
    MiniMax, xAI Grok, Writer Palmyra, AI21 Jamba, TwelveLabs Marengo video embeddings, and others.

This is a hand-picked sample, not the full roster — the Models page lists every model this gateway actually serves, generated from the live catalogue.

Model Management

  • Automatic model discovery — Configured regions are scanned at startup, so there is no model list to maintain by hand and nothing to keep in step with Bedrock's catalogue
  • Aliases — A model is published under whichever names you choose, with Claude and OpenAI names resolving on their own; an alias can carry its own service tier, guardrail, metadata and parameters, so one model serves several policies under several names
  • Wildcard model names — Any request that names a model can name a glob pattern instead (claude-sonnet-*), and the newest release that matches serves the request — resolved once, at creation, so a pinned job keeps its model for its whole life
  • Deprecation handled for you — Models AWS has retired drop out of the list your users pick from, so nobody builds on a model about to be withdrawn, and requests naming one are redirected to its replacement; a workload that still depends on one can keep it listed
  • Capability discovery — The catalogue advertises what each model can actually do — modality, route, streaming, speech to speech, transcription and translation, the search surfaces, and Batch API support — filterable over HTTP and through the same tool an agent reads before it calls anything
  • Published, not just discoverable — The Models page lists every model this gateway serves on AWS, with modalities, regional availability, AWS prices and independent leaderboard scores

Multi-Modal Capabilities

Text & Conversational AI

  • Server-side conversations — a thread is kept server-side and continued by id instead of resending the history, its items listed and managed, and a response attached to it through the Responses API conversation parameter; a long thread can be compacted into a reusable summary item rather than replayed in full, on request or automatically past a token threshold, or have its oldest turns dropped once it outgrows the model's context window
  • Token counting before a call, so a prompt can be sized against a model's window without paying to generate
  • Streaming over Server-Sent Events with tokens delivered as they arrive; reasoning content blocks on the models that produce them, and web search results as context
  • Image, document, audio and video attachments on multimodal models — see Attachment Size for how large ones are carried

Images

  • Generation — Text-to-image in PNG, JPEG or WebP, at the size, aspect ratio, quality and compression asked for, with model-specific style presets and partial previews streamed while it renders
  • Editing — Mask-based inpainting over a region you define, image-to-image transformation on style or structure, background removal, object search-and-replace, object recolor, and creative or conservative upscaling
  • Variations — Alternative versions of an existing image
  • Nothing re-uploaded — An input image is referenced by Files API file_id or by URL rather than sent again with every call

Audio

Text-to-speech (Amazon Polly)

  • 60+ voices across 30+ languages, on the Standard, Neural, Long-Form and Generative engines, with the language detected automatically
  • SSML control over pronunciation, emphasis, pauses and prosody; 0.2× to 4.0× speed on plain text, set in the document itself for SSML input; MP3, PCM, Opus, AAC, FLAC and OGG Vorbis output
  • Long input — up to 100,000 characters per request, 24× OpenAI's limit (20,000 with a generative voice, which speaks it as the audio is delivered)

Speech-to-text (Amazon Transcribe)

  • 100+ languages, detected automatically when the request does not name one
  • Speaker diarization, word-level and segment-level timestamps, and SRT/VTT subtitle export
  • Vocabulary customization and custom language models, per language — so a request identifying between several applies the right resources to each one
  • Streamed results — each phrase comes back as it is recognized instead of after the whole recording, whenever the request names the language to expect
  • Transcripts encrypted with your own KMS key, under a key policy scoped to this workload rather than to a whole bucket
  • Medical transcription — clinical dictation and patient–clinician conversations in US English, with six medical specialties when streamed

Speech translation — Transcribe audio and translate to English in a single request; a language pair that cannot be served is refused as a request problem instead of failing after the audio has been transcribed.

Speech-to-text (Amazon Nova Sonic) — An alternative backend on both audio routes, at the lowest transcription cost available here — about $0.006 per minute of audio at current Amazon Bedrock rates. Punctuated transcripts in the language spoken, with automatic language detection, and translation to English produced by the model itself in one request. json and text output only, up to 10 minutes of audio per request — no timestamps, subtitles, diarization or detected-language reporting.

Live speech-to-speech (Realtime API)

  • Bidirectional audio over a single WebSocket, OpenAI Realtime API compatible
  • 24 kHz PCM, or G.711 at 8 kHz for telephony; server-side voice activity detection or manual turn control, with barge-in on the item the caller spoke over
  • Ephemeral, browser-safe client secrets, minted by one instance behind a load balancer and verified by any other
  • Function calling: the model asks for one of the session's own functions, the application runs it, and the answer is spoken with the result
  • A configured guardrail is applied per turn — a written item is checked before it reaches the model, a spoken answer once it is complete (guardrail coverage)
  • Its limits, up front — a session lasts at most 8 minutes. WebSocket is the default transport; WebRTC calls terminated by the gateway are an operator opt-in with documented single-instance and UDP-ingress constraints, and SIP is never terminated in-process. For telephony or multi-instance WebRTC, LiveKit Agents and Pipecat terminate the media themselves and reach this API like any other client — see transports and the feature compatibility table

Documents & Files

  • PDF input with optional citations — the answer points back at the exact source passage
  • Plain text and structured content blocks as context; a large PDF or document is carried by reference on a model that reads it from storage — see Attachment Size
  • Upload once and reference by ID across requests, with expiry anywhere from 1 hour to 30 days
  • Multipart uploads for large files, backed by S3's own multipart

Video

  • Text-to-video and image-to-video generation (Amazon Nova Reel, Luma Ray 2) through the OpenAI Videos API, as an asynchronous job — create, list, poll, download, delete
  • Video as chat input on the models that read it (Amazon Nova among them); long clips follow the Attachment Size policy
  • s3:// URLs as direct video input for multimodal embeddings

Attachment Size

On chat completions, messages and responses served by Amazon Bedrock, every attachment — an image, document, audio or video sent as base64, a data URI, an HTTPS URL, an s3:// URI or a Files API ID — is measured before the request is built, and carried the way that model accepts:

Attachment size How it is sent
Within the model's inline capacity Embedded in the request
Above it, on a model that reads that kind from storage Staged in a bucket of yours in the region serving the request, and referenced
Above it, on a model that reads it inline only Refused with 413, stating the size that model accepts

Models differ in their per-attachment and per-request limits and in which kinds they read from storage — the Amazon Nova families read images, documents and video that way and TwelveLabs Pegasus reads video, while the rest read attachments from inside the request — so the same file can be inline for one model and staged for another, with nothing in the request changing either way. An attachment already in S3 is referenced as it stands whatever its size, on a model that reads that kind from storage. Bedrock Mantle-served models, image editing and variations, transcription and the embeddings routes keep their own input handling.

Embeddings

  • Text embeddings, single or batched, and multimodal embeddings over images, audio, video and PDF documents
  • Dimension reduction where the model offers it, in float or Base64 encoding
  • s3:// input for large files, with oversized base64 payloads staged to S3 for you

Retrieval & Vector Stores

There is no embedding pipeline, chunker or vector database to run alongside the gateway. The Vector Stores API takes an attached file, chunks it, embeds it and indexes it in the background, reports the indexing as it progresses, then searches it by meaning.

  • Search by meaning, with the passages — Each result carries its file, its text, its score and the attributes stored with it
  • Attribute filters — Tag a file with up to 16 attributes and restrict a search to the ones that match
  • File batches and expiration policies — Attach many files under a single identifier; expire a store after a number of days without a search
  • Held in your own account — Documents, passages and vectors live in an Amazon S3 vector bucket in your account, not in an index somebody else runs

Point it at the knowledge base you already run

A store can equally be an Amazon Bedrock knowledge base you already run, addressed through the same endpoints — searched, its documents attached, listed and read. Bedrock-managed and customer-managed knowledge bases over documents are both served, whichever way you built yours; one built over a structured data store or an Amazon Kendra index is not, since it answers with database rows or search hits rather than passages.

It stays yours. Knowledge bases are allowlisted one by one, never created or deleted here, and anything that would reshape one — renaming, expiry, chunking strategy, rewriting a file's attributes — is refused, naming why. A request naming a knowledge base that is not on the allowlist answers exactly as one that does not exist, so the allowlist cannot be probed for what a deployment holds.

The model does the retrieving

file_search on /v1/responses gives any chat model served on that route the stores you name, whether managed or knowledge-base backed. The model decides when to search and with which query; each search is reported as a file_search_call item carrying the queries it used (and the passages themselves on request, plain or streamed), and the grounded answer carries a file_citation annotation for every file it drew on. A filter operator the serving store cannot apply, or a score threshold against a store whose scores have no defined scale, is refused with a 400 rather than quietly dropped — so an answer does not come back as though it had honoured a restriction it ignored.


Purpose-Built for AWS

Multi-Region Routing & Quota Headroom

A deployment spanning several AWS regions draws on more than one Bedrock quota and keeps serving when one region is degraded. How traffic spreads across them is yours to choose, and the choice trades throughput against prompt-cache hit rate:

Routing across regions What it gets you Prompt Caching
In order Deterministic placement; blocked regions skipped ✓ Compatible
Lowest latency The fastest measured region for each call ✓ Compatible
Round robin Load spread evenly, at the cost of cache locality not compatible
Single region Every call to a model served from one place ✓ Compatible
  • Each region adds its own quota — Bedrock tokens-per-minute and requests-per-minute limits are per region, so a multi-region deployment draws on several independent quotas rather than one. How much of that headroom a workload reaches depends on the quota granted per model in each region and on the routing strategy
  • Eligible failures retry elsewhere — A throttle, quota or service error switches region transparently, under a backoff that widens while a region keeps failing. Streaming responses can only retry before the stream opens, and asynchronous jobs stay in the region that accepted them
  • Health is tracked per model — A region that failed for one model is set aside for that model alone and brought back once it recovers, rather than taking the whole catalogue down with it

Resilience & Failover

Advanced Bedrock Features

Feature Description
Prompt Caching Cache system prompts, messages and tools section by section, at the TTL you choose — a long system prompt is billed at the cache-read rate on later turns instead of in full, with cache metrics in the response; automatic, with no parameter, on Moonshot Kimi K3 and OpenAI GPT-6
Reasoning Modes Extended thinking on Claude, Nova, OpenAI GPT, Kimi and DeepSeek, driven by effort level or by a token budget
Context Editing Claude clears the oldest tool results and thinking blocks of a long agent conversation before it reaches the model — Anthropic's context_management, with the applied edits reported back and honored by token counting
Bedrock Guardrails Content filtering and safety policies applied to traffic from every client, with the trace detail you choose
Service Tiers Priority, default, flex and reserved tiers, per request or as a default per model — a latency-sensitive workload and a cheap bulk one share the deployment
Application Inference Profiles Isolate a workload and see it separately on the AWS bill
Prompt Routers Bedrock prompt routers for intelligent model selection
Cross-Region Inference Geography-pinned (US, EU, APAC) and global profiles, so inference stays inside the geography your data residency requires
Web Search / Grounding Built-in web search with source citations, billed per query: Amazon Nova grounding (Chat Completions, Responses, and Messages) and OpenAI GPT built-in search (/v1/responses only). A deployment can keep grounding off the open internet entirely, and a request asking for what it forbids is refused rather than silently rewritten
Server-side tools Amazon Nova's code interpreter, and Claude's bash, text editor, computer use and memory tools on the generations that carry them

Asynchronous Batch Inference

Large request sets run asynchronously at Amazon Bedrock's discounted batch price, on both dialects — /v1/batches and the Anthropic /v1/messages/batches. Submit, poll, collect, cancel.

  • Chat or a whole corpus — On the OpenAI surface a batch is a JSONL file of chat completion or embeddings requests, each result carrying its own custom_id
  • A model per request — An Anthropic batch may name a different model for each request and is still submitted, tracked and collected as a single batch, whatever it fans out to
  • Priced as batch — Usage is recorded and priced at the tier that actually served the call, so a batched request is reported at the batch rate rather than the on-demand one
  • Discoverable before you submit — The model catalogue reports and filters on Batch API support. It is best effort: a model without the flag is still submitted, since the absence may only mean no price is published for it yet
  • Result files expire on your terms — rather than being kept until deleted

Batches run under a role of yours

Amazon Bedrock reads the requests and writes the results itself, under an IAM role and a bucket of yours — the batch runs on your account's terms, not the gateway's. See Batch inference IAM.

Bedrock Mantle Models

stdapi.ai serves models from the Amazon Bedrock Mantle endpoint alongside the classic Bedrock catalog: OpenAI GPT, xAI Grok, Google Gemma, Qwen, GLM, DeepSeek, MiniMax, Kimi, Nemotron, and more — the available catalog varies per region and grows over time.

  • Every text API, every model — All four text APIs (chat completions, responses, messages, legacy completions) work with every Mantle model: served natively when the model supports the API upstream, converted automatically otherwise
  • Routing you choose — A model available on both endpoints is served by the classic one, except the OpenAI GPT-5.6 and GPT-6 families, which are served by Mantle where it lists them so that their built-in web search and code interpreter work without configuration; Mantle serves the models only it has. Any dual-homed model can be pointed at Mantle — or taken back off it — for the whole deployment, or routed for a single request, tapping Mantle's separate throughput quotas on top of your Bedrock ones
  • The same operational behaviour — Region failover, quota backoff, usage recording and pricing work as they do on classic Bedrock; requests chained via previous_response_id stay pinned to their origin region; and access runs on the same AWS credential chain, with no separate API key to issue, store or rotate
  • Native stored conversations — /v1/responses with store and previous_response_id uses Mantle's own server-side storage: 30-day retention, region-local, project-scoped
  • Built-in web search — The OpenAI GPT-5.x and GPT-6 families ground answers in current web content with source citations on /v1/responses, inside the AWS boundary by default

Mantle models appear in the same catalogue as the rest, under the same /v1/models call — nothing in a client distinguishes them. A deployment whose IAM policy does not reach Mantle simply does not list them.

Price change for the OpenAI GPT-5.6 and GPT-6 families

Serving GPT-5.6 Sol, Terra and Luna and GPT-6 Astra on Mantle is a price change, not only a routing one. Both endpoints charge the same In-Region rate, but Mantle has no cross-region inference profiles, so these models no longer ride the Global profile and its discount: every token costs exactly 10% more than on the classic endpoint's default Global routing — input, output, cached and long-context rates alike. The per-million figures are in the AWS_BEDROCK_MANTLE_PREFERRED_MODELS reference. A deployment that pins In-Region routing pays what it paid. GPT-6 Sol and Luna have no published rate yet, so their usage is recorded without a cost.

Three other things change with them: Amazon Bedrock Guardrails cannot apply (a deployment configuring both is refused at startup), and usage and cost are reported — and billed by AWS — under Bedrock Mantle rather than Bedrock, attributed by project. Throughput runs on Mantle's own quotas. Batch inference, prompt caching and stored responses are unaffected: batches still run on the classic endpoint, both endpoints cache, and response IDs issued before the change keep working.

Set AWS_BEDROCK_MANTLE_PREFERRED_MODELS to an empty value to serve every dual-homed model on the classic endpoint, at the classic price and under your guardrail.

Limitations & conversion details

Bedrock Guardrails and cross-region inference profiles do not apply to Mantle-served requests, and the built-in web_search tool is served on /v1/responses only. API-shape conversion preserves the core request semantics (messages, tools, sampling, streaming, usage); parameters with no equivalent in the serving API are dropped or adapted. The exact parameter tables, response-ID specifics, and per-route limitations are on the API pages: chat completions, responses, messages, and legacy completions.

Bedrock Mantle Configuration

Bedrock Marketplace Model Endpoints

Deploy a model from the Amazon Bedrock Marketplace catalog onto a managed endpoint in your own account, and stdapi.ai serves it beside the rest — same /v1/models listing, same chat completions, responses and messages APIs, no client-side change.

  • Discovered, never deployed — The gateway lists the endpoints you deployed and publishes the ones that are ready to serve. It never creates, updates or deletes one: that is capacity you pay for, and it belongs in your own infrastructure-as-code
  • Opt-in and region-scoped — Off unless you enable it, and only in the regions you serve: an endpoint is invoked in its own region, so one deployed elsewhere is never published

These endpoints bill by the hour, not by the token

A model endpoint runs on dedicated instances and is charged for every hour it exists, whether or not anything calls it, and there is no scale-to-zero on this path. stdapi.ai reports the tokens it served and no cost, because AWS publishes no per-token rate for them — your bill is instance-hours. See Marketplace model endpoint costs.

Marketplace Model Endpoints Configuration

SageMaker AI Endpoints

Run your own model — a fine-tune, an open-weight release, anything the SageMaker AI vLLM or SGLang containers can serve — on an Amazon SageMaker AI endpoint in your account, and stdapi.ai serves it beside the rest: same /v1/models listing, same chat completions, responses and messages APIs, no client-side change.

  • Named, never deployed — You name the endpoints you want served, with the model ID your clients should ask for. The gateway invokes them and nothing else: it never creates, updates, scales or deletes an endpoint, because that is capacity you pay for and infrastructure you own
  • A cold start your callers never see — An endpoint scaled to zero has no capacity to answer with, and the request itself is what makes AWS provision an instance again. stdapi.ai holds the connection and retries until the model answers, so the caller gets a slow first request rather than an error. Concurrent callers share one wait
  • Every AWS partition — SageMaker AI endpoints exist in the commercial, GovCloud, China and European Sovereign Cloud partitions alike, which makes this the way to serve a chat model where no other backend reaches

These endpoints bill by the hour, not by the token

A SageMaker AI endpoint is charged by the instance-hour for as long as it has instances running, whether or not anything calls it. stdapi.ai reports the tokens it served and no cost, because AWS publishes no per-token rate for them — your bill is instance-hours. An endpoint configured to scale to zero costs nothing while idle, at the price of that first slow request. See SageMaker AI endpoint costs.

What the container decides

The endpoint's container decides what the model can do: tool calling needs a tool-call parser configured on it, reasoning content needs a reasoning parser, and image input needs a model and a container that accept image content parts. The gateway publishes what you declare and forwards what you send; it cannot add a capability the container does not serve. Token counting answers an approximation for these models that errs high — see input token counting.

Guardrails do not apply

A container serves the OpenAI Chat Completions API and carries no guardrailConfig, so an Amazon Bedrock Guardrail cannot filter what these models answer. A request that reaches one while a guardrail is configured is refused with a 400 rather than served unfiltered, and a guardrail-bearing model alias naming one of them stops the server at startup.

SageMaker AI Endpoints Configuration

Amazon S3 as the file layer

S3 backs the whole API surface, not just file storage, which buys three things a file API bolted onto a database cannot:

  • No artificial size ceiling — Files reach ~78 GiB in one direct upload and 48.8 TiB per upload session, in native multipart parts, streamed rather than buffered. One file ID works on both the OpenAI and the Anthropic endpoints
  • s3:// is a first-class input — An object already in your buckets is named directly in chat completions, Messages, embeddings and image operations, read under the gateway's IAM role: no pre-signed URLs, no download-and-re-upload round trip
  • Region-local by construction — Anything the gateway stages sits in a bucket in the region serving the request, so payloads do not cross a region on the way to the model; a generated image can be handed back over S3 Transfer Acceleration, downloaded from a CloudFront edge instead of the bucket's region

Security & Compliance

Authentication

Method How Best For
API Key Authorization: Bearer or X-API-Key header; stored in SSM Parameter Store or Secrets Manager (never plain text) Direct clients, SDKs
Cognito user pool JWT Authorization: Bearer with an Amazon Cognito access token, validated per request Per-user access, agents
OIDC / Cognito Delegate to AWS Application Load Balancer or API Gateway Web apps, SSO
AWS IAM (SigV4) Via API Gateway with IAM authorization Internal AWS services
No authentication Open access Private VPC deployments
  • Per-caller identity — Amazon Cognito user pool tokens are accepted instead of, or alongside, the API key, so each caller reaches the API with their own credential — validated in-process against the pool's published keys, with no AWS call on the request path, and it is that verified identity per-user cost attribution bills against
  • The posture is asserted, not inferred — Name the method you intend to run: the server refuses to start when the method you named is not actually in force, or when one you configured would be silently ignored, so a deployment cannot drift into answering unauthenticated traffic
  • Agents authenticate without being configured — Every unauthorized response points at the document naming the authorization server and the scope this deployment expects

Authentication & Security

Security Features

  • A URL in a request stays outside your network — Loopback, link-local and private addresses are refused, along with DNS rebinding; hostname allowlisting, CORS policy and CSRF protection govern what may call the service and from where
  • Malformed requests stop at the edge — Out-of-spec requests are rejected before they reach an AWS call, under a strict mode that also refuses unknown fields instead of ignoring them
  • API keys are not stored in the clear — Held in SSM Parameter Store or Secrets Manager, kept in memory only as a salted hash, and compared in constant time so a key cannot be recovered by timing
  • Encrypted in transit — TLS 1.2+ on every AWS service call; the Terraform module terminates client traffic on TLS 1.3 with post-quantum hybrid key exchange, and forwarded headers from ALB and CloudFront are processed safely
  • A hardened supply chain — The commercial container image is validated and built without exposure to a public package registry at run time

Commercial: hardened image, Security Hub validated AWS Marketplace

The commercial image is security-validated by AWS Marketplace: a hardened minimal base with minimal installed packages and no shell, deployed with a read-only root filesystem and dropped Linux capabilities. The community image is a standard Debian-slim build carrying the same API surface, the same models and the same media stack, so what the commercial edition adds is hardening and optimization, not capability. The Terraform module is built against the AWS Security Hub Foundational Security Best Practices standard, passes a large share of applicable controls out of the box, and configures a Customer Managed KMS key with auto-rotation for all data at rest. Optional variables add native GuardDuty Runtime Monitoring and Route 53 Resolver DNS Firewall on the module's dedicated VPC.

AWS Security Hub, GuardDuty & DNS Firewall Integration

Compliance & Data Sovereignty

The gateway runs on infrastructure you own, so no third party sits between your users and your models, and Amazon Bedrock does not store your prompts or use them to train models. AWS service calls are restricted to the regions you configure. The AWS services used by stdapi.ai (Bedrock, S3, Polly, Transcribe, and more) are in scope for GDPR, ISO 27001/27017/27018, SOC 1/2/3, HIPAA, FedRAMP, PCI-DSS, and CSA STAR Level 2 — these certifications apply to the AWS services and regions you choose, and are not inherited by stdapi.ai or by your application. The commercial Terraform module adds VPC endpoints (no internet egress), Customer Managed KMS keys, and region-pinned cross-region profiles for strict data residency.

Data Sovereignty & Compliance


Works with Your Existing Tools

stdapi.ai speaks the APIs hundreds of applications and tools already speak, so adoption is quick: an application points at your deployment instead of the vendor's, with the key your gateway issues. Model names carry over — Claude and OpenAI names resolve on their own — and the name is now drawn from every provider in the catalogue rather than one vendor's list. A name the catalogue does not hold returns 404 instead of a lookalike, and a served model can be published under whichever name your application already sends, so an application whose model name is not yours to change keeps working untouched.

  • Chat Interfaces
    Open WebUI, LobeHub, AnythingLLM, LibreChat — private ChatGPT-style experiences on AWS

  • AI Coding Assistants
    Claude Code, Cline, OpenCode, Pi Agent, Zed — backed by Claude, Kimi, Qwen3 Coder

  • Workflow Automation
    n8n, Langflow, Dify, Flowise — connect AI to your business processes

  • Agent Frameworks
    OpenClaw, Hermes Agent, LangChain, LangGraph, CrewAI, OpenAI Agents SDK, Pydantic AI, Agno, Strands Agents — multi-agent systems on Bedrock

  • Voice & Audio
    Pipecat, LiveKit Agents, TEN Framework, Home Assistant — voice agents on live speech-to-speech, transcription, and translation

  • RAG & Semantic Search
    LlamaIndex, Haystack, RAGFlow, Docling, LightRAG — built-in vector stores, embeddings and Cohere-compatible reranking

Team chatbots in Slack, Discord or Microsoft Teams and knowledge tools such as Obsidian Copilot, Khoj and SiYuan connect the same way.

See all use cases


AI Agents

Agent Discovery

An agent that has only the base URL can work out the rest for itself, through standards rather than a hand-written integration: RFC 8288 Link headers on / and an RFC 9727 catalog at /.well-known/api-catalog point at the OpenAPI schema, the documentation and the SEP-1649 MCP server card, which advertises the transports on offer. It can work out how to authenticate itself the same way: an RFC 9728 protected resource metadata document names the authorization servers issuing tokens for this deployment and the scopes they need, and every 401 carries its address in the WWW-Authenticate challenge — so an MCP client reaches a secured deployment it was never configured for.

MCP (Model Context Protocol)

stdapi.ai exposes its API surface as MCP tools, letting AI agents and orchestrators call an endpoint directly through the Model Context Protocol — no HTTP client code required. Every operation is published except the ones an agent could not use: the Realtime WebRTC call verbs, which need a peer connection it has no way to hold, and the Ollama model-store verbs that always refuse. The organization usage and costs tools appear only where the settings behind them — the usage API, CloudWatch metrics and cost tracking — are enabled, since they would otherwise answer 503. Naming any of them in the include list publishes it anyway.

  • 90+ tools, no integration code — Each published API operation (chat, images, audio, embeddings, files, model search) is a named MCP tool with generated documentation, over Streamable HTTP at /mcp or SSE at /sse for older clients
  • A tool surface you choose — Tools are included or excluded by name, so an agent is handed exactly the capabilities it should have and no more — a read-only deployment, or one without file deletion, is a list away
  • Written for an agent's context window, not a human's — Schemas hide parameters an MCP client cannot use and results come back as compact JSON, so each call costs the calling agent fewer tokens
  • Media-aware results — An endpoint answering with bytes returns an image or audio result the agent can use directly, falling back to a download reference for what the protocol cannot carry, such as a generated video

MCP Configuration


Observability & Operations

Logging & Tracing

  • JSON to stdout, ingested by CloudWatch as it stands — Every request logs its method, path, status, model, the region or regions that served it and how long it took, so a slow model or a region that started failing is one query away
  • Prompts stay out of your logs unless you ask for them — Full request and response payloads and the client IP are available when you are debugging and are not written otherwise
  • Traces and metrics into what you already run — OpenTelemetry export to AWS X-Ray, Datadog, Jaeger or any OTLP backend, one root span per request with correlation IDs, sampled at a rate you set

Cost Tracking

  • Usage counts read back from AWS — Token, character, second and image counts come from the AWS responses themselves rather than from client-side counting; recorded per request across chat, embeddings, images, audio and built-in tools, and reported to the caller in the same usage shape on every endpoint — input, output, reasoning and cached tokens included
  • Priced on the dimensions AWS bills on — From AWS's own price list, refreshed automatically with no list to maintain by hand: serving region, service tier (standard, flex, priority or batch — the tier that actually served the call), prompt-cache TTLs, cross-region and latency-optimized routing, long-context rates, image resolution and quality. Operator overrides cover any gap
  • Beyond tokens — Built-in web searches are counted per query, and a search against a knowledge base the backend manages at its published rate per call. What cannot be accounted for is stated rather than approximated
  • Currency-safe figures — Cost appears in the request log in your AWS partition's own currency (USD, EUR, CNY), as exact decimal amounts that are never summed across two
  • Model Pricing API — The loaded catalog is queryable at GET /model_pricing, over HTTP or as an MCP tool, for cost-aware model selection; spend can also be published to CloudWatch as EMF metrics
  • Usage and cost reporting in the shape your tools already read — Consumption and spend come back from the OpenAI Administration usage and costs endpoints, answered from the metrics the deployment publishes to Amazon CloudWatch. It reports the whole deployment, so it is an administrator surface: off by default, and never readable with a tenant key

An estimate, not a bill

Costs are estimated from AWS's published prices, not read back from your invoice — a best-effort figure for visibility and alerting. See Cost Tracking for its accuracy and known limitations.

Per-User Cost Attribution

  • Each end user on the AWS bill — Model calls run under a short-lived role session opened for the user behind the request, so AWS reports their spend separately in Cost Explorer and CUR 2.0 — from the invoice, not from an estimate
  • The identity the gateway verified — The authenticated caller where authentication is enabled, otherwise the identifier the request declares (safety_identifier/user, or metadata.user_id on the Anthropic Messages API). It travels as a session tag: a cost allocation dimension in Cost Explorer, and an access boundary testable in IAM as aws:PrincipalTag
  • Fail-closed — A session that cannot be opened fails the request rather than quietly billing the gateway, and requests identifying no user can be rejected outright

It covers model invocations

Each user's model calls are attributed; the rest of the gateway's AWS usage stays on its own identity — see Per-User Attribution for the role those calls run under.

Day-to-Day Operation

  • An API reference on the deployment itself — Swagger UI at /docs to try an endpoint in a browser, ReDoc at /redoc to read it, and the OpenAPI schema at /openapi.json to generate a client or import into Postman
  • Proxy-aware outbound connections — HTTPS_PROXY, HTTP_PROXY and NO_PROXY are honoured by the connections the server makes to AWS and to model endpoints, not by the AWS SDK alone (proxied deployments)
  • A missing permission reads as one — An unconfigured resource or a denied AWS call answers 503 feature_unavailable, with the server log naming the operation, the model and the permission AWS refused, instead of reaching clients as their own key being rejected

Performance

A gateway earns its place by adding as little as possible on top of the model call. Measured gateway CPU on the production serving stack, single worker, over the complete request path:

Request shape Gateway CPU per request
Typical chat request (2.5 KB) 0.8 ms
Large context (1 MB body) 4.6 ms
Large context, streamed (~100 events) 8.6 ms

Independent work fans out concurrently; JSON, the AWS wire format and the HTTP serving stack all run compiled; and a streamed response is passed through as it arrives rather than buffered — which is why those figures hold precisely where load does, on large contexts and streaming.

Negligible next to the model call

Even at its most expensive — a 1 MB request — the gateway's processing adds a few milliseconds to an invocation the model itself takes seconds to answer: well under 1% of end-to-end latency. Measured live, a typical chat completion spends about a millisecond in the gateway out of a several-hundred-millisecond round trip — a share that holds even with the server capped to 0.25 vCPU, the smallest Fargate task size.


Quality Assurance

Every compatibility claim on this page is tested against the real vendor APIs and driven by real client software — see the evidence.


Deployment

Community vs Commercial

Community Commercial
Price Free $0.10/container-hour - With 14-day free trial
License AGPL-3.0 AWS Marketplace SCMP
API compatibility Full Full
Container image Standard Debian-slim build (GHCR) Hardened minimal base, AWS Marketplace security-validated
Deployment Docker / self-managed Terraform module (ECS Fargate) - AWS Marketplace container image
Production infrastructure not available Fully featured - AWS Well-Architected - Hardened
Security posture Manual (self-managed) Built against Security Hub FSBP, passing most applicable controls out of the box; GuardDuty & DNS Firewall integrations
Commercial support not available 1 business day

How stdapi.ai Compares

Feature by feature against the other ways to put an OpenAI-compatible API in front of Amazon Bedrock, with sources and verification dates — see the comparison.


Next: Get Started · Deploy on AWS · Run Locally · API Reference · Models