Chat Completions API¶
Generate conversational AI responses with Amazon Bedrock foundation models—including Claude, Nova, Llama, and more—through an OpenAI-compatible interface.
At a glance¶
- Every Amazon Bedrock chat model behind one schema — Anthropic Claude, Amazon Nova, Meta Llama and the rest of the Converse catalogue, plus the Bedrock Mantle models (OpenAI GPT-5.x, xAI Grok, Google Gemma 4) where it is enabled.
- Multi-modal messages. Images, video, documents and audio alongside text, referenced by HTTPS URL, data URI,
s3://URI, or afile_idfrom the Files API. - Reasoning in the dialect your client speaks.
reasoning_effort(OpenAI),enable_thinkingandthinking_budget(Qwen),thinking(Moonshot) and the OpenRouterreasoningobject all reach the same control. - Prompt caching and stored completions. Cache breakpoints map to Amazon Bedrock
cachePoints, andstore: truekeeps the completion in Amazon Bedrock session storage in your own AWS account. - Content safety per request. An Amazon Bedrock guardrail applies through a request header or the
moderationparameter, on requests served by the Converse API. - Web search where the model runs one.
web_search_optionsgrounds the answer in a live search and cites it in the messageannotations— see Web Grounding. - Differs from OpenAI:
logprobsis rejected with400whiletop_logprobskeeps working,n > 1is not available with streaming, andweb_search_optionsneeds a model that runs a web search — it is rejected elsewhere rather than answered ungrounded.
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "amazon.nova-micro-v1:0",
"messages": [{"role": "user", "content": "Say hello world"}]
}'
Endpoints¶
| Endpoint | Method | What It Does | Powered By | MCP Tool |
|---|---|---|---|---|
/v1/chat/completions | POST | Conversational AI with multi-modal support | Amazon Bedrock Converse API · Amazon Bedrock Mantle | openai_chat_completion |
/v1/chat/completions | GET | List stored chat completions | Amazon Bedrock Sessions | openai_chat_completion_list |
/v1/chat/completions/{completion_id} | GET | Retrieve a stored chat completion | Amazon Bedrock Sessions | openai_chat_completion_get |
/v1/chat/completions/{completion_id} | POST | Update a stored chat completion's metadata | Amazon Bedrock Sessions | openai_chat_completion_update |
/v1/chat/completions/{completion_id} | DELETE | Delete a stored chat completion | Amazon Bedrock Sessions | openai_chat_completion_delete |
/v1/chat/completions/{completion_id}/messages | GET | List the input messages of a stored chat completion | Amazon Bedrock Sessions | openai_chat_completion_messages |
Feature compatibility¶
Two outcomes are possible for a parameter no model behind this API can honor, and the Notes column below says which one applies. A parameter that only tunes an answer you can still use — a quality or latency hint, a telemetry opt-in — is accepted and ignored, so a client setting it on every request keeps working. A parameter that is the request — an output format, a modality, or a safety restriction that would silently disappear — is rejected with a 400 naming it, because the alternative is returning something you did not ask for.
| Feature | Status | Notes |
|---|---|---|
| Messages & Roles | ||
| Text messages | Full support for all text content | |
Image input (image_url) | HTTP, data URIs | |
| Image input from S3 | S3 URLs | |
| Video input | Supported by select models | |
| Audio input | Supported by select models | |
Document input (file) | PDF and document support varies by model | |
Assistant audio reference | An assistant turn carrying only audio: {"id": …} is dropped (past audio is not replayable); resend the transcript as text to keep it in context | |
Message name (participant) | Names a speaker among same-role turns. Accepted and ignored — the Amazon Bedrock Converse Message has no author field, so name each speaker inside the message text where it has to reach the model; forwarded upstream on Mantle passthrough models | |
Files API (file_id) | Reference uploaded files via type: "file" — see Files API | |
| System messages | Includes developer role | |
| Tool Calling | ||
Function calling (tools) | Full OpenAI-compatible schema | |
Legacy function_call | Backward compatibility maintained | |
| Parallel tool calls | Multiple tools in one turn | |
| Disable parallel tool calls | parallel_tool_calls: false is accepted for every model and honored by models able to constrain tool use; the response reports the tool calls actually made | |
| Server tools | Provider system tools and Claude server tools | |
tool_choice | auto, none, required, and named-function choice are supported; none withdraws every declared tool, but a conversation already containing a tool_calls message keeps the tools it names callable; tool_choice: {"type": "allowed_tools"} is rejected with 400 — supported on the Responses API; Claude Opus and Sonnet 5.5 and later refuse required, a named function and a named legacy function_call with 400 | |
Changing tools between turns | The tools declared on a turn are the only ones the model may call, whatever the replayed history names | |
custom tools (free-form) | Rejected with 400, in tools, in a named tool_choice, and inside allowed_tools; declare the tool as a function tool instead — accepted and dropped on the Responses API | |
| Generation Control | ||
max_tokens / max_completion_tokens | Output length limits | |
temperature | Mapped to Bedrock inference params. The full 0–2 range is accepted; a value above 1.0 is served at 1.0 (an Amazon Bedrock limitation) | |
top_p | Nucleus sampling control | |
stop sequences | Custom stop strings. Generation halts at the first sequence matched and nothing after it is returned. Unlike the OpenAI API, which always strips it, whether the matched sequence is returned at the end of the text is model-dependent — do not rely on it as an end-of-answer marker. Some models refuse the parameter with 400 rather than ignoring it. Whitespace-only sequences are rejected with 400 (Amazon Bedrock limitation) | |
frequency_penalty / presence_penalty | Repetition control | |
seed | Deterministic generation | |
logit_bias | Not all models support biasing | |
top_logprobs | Forwarded to the model as a provider-specific field; honored only by models that support it. Usable even though logprobs is rejected | |
top_k (From Qwen API) | Candidate token set size for sampling | |
reasoning_effort (OpenAI API-compatible) | Reasoning control: none/minimal/low/medium/high/xhigh/max (accepted for all models) | |
enable_thinking (Qwen API-compatible) | Enable/disable thinking mode (accepted for all reasoning models) | |
thinking_budget (Qwen API-compatible) | Thinking token budget (accepted for all reasoning models) | |
thinking (Moonshot API-compatible) | Thinking config: {"type": "enabled"/"disabled"} (accepted for all models) | |
reasoning (OpenRouter API-compatible) | Reasoning object: effort, max_tokens, enabled, exclude. Equivalent to reasoning_effort, thinking_budget and enable_thinking; conflicting values are rejected with 400 | |
include_reasoning (OpenRouter API-compatible) | false omits the reasoning text from the response, like reasoning: {"exclude": true}; the reasoning tokens are still generated and billed | |
n (multiple choices) | Generate multiple responses, not supported with streaming | |
logprobs | Rejected with 400 when enabled (false/null accepted, as they request the default behavior); top_logprobs (above) remains usable | |
prediction | Static predicted output content. Accepted and ignored — a latency hint the Amazon Bedrock Converse API has no equivalent for, and the answer is unchanged without it; forwarded upstream on Mantle passthrough models | |
response_format: "json_object" | Accepted for all models; syntactically valid JSON is not guaranteed for every model, and a model may wrap it in a Markdown code fence. Use json_schema where the output has to be machine-readable without post-processing | |
response_format: "json_schema" | Structured JSON output validated against the supplied schema | |
verbosity | Model verbosity. Accepted and ignored — steer the answer length from the prompt; forwarded upstream on Mantle passthrough models | |
web_search_options | Grounds the answer in a web search, cited in the message annotations — see Web Grounding. Rejected with 400 on a model that runs no web search. search_context_size is accepted and ignored; user_location is rejected with 400 | |
translation_options (Qwen API-compatible) | Translation tuning options. Rejected with 400 when set | |
prompt_cache_key | Cache prompts to reduce costs and latency | |
prompt_cache_options | mode: "explicit" caches only the parts marked with prompt_cache_breakpoint; ttl: "30m" mapped to a 1 hour Amazon Bedrock retention on Anthropic models (other models use the default 5 minute TTL) when prompt_cache_retention is unset | |
prompt_cache_breakpoint (content part) | Explicit cache boundary mapped to an Amazon Bedrock cachePoint (max. 4 per request) | |
| Extra model-specific params | Extra model-specific parameters not supported by the OpenAI API | |
| Streaming & Output | ||
| Text | Text messages | |
Streaming (stream: true) | Server-Sent Events (SSE) | |
| Streaming obfuscation | stream_options.include_obfuscation is accepted and ignored — the stream is never padded | |
| Audio | Model output or synthesis from text output (synthesis is Converse-only — not performed for Mantle-served requests); non-streaming only — stream: true with audio output is rejected with 400 | |
response_format (JSON mode) | json_object accepted for all models, without a syntax guarantee on every model and possibly wrapped in a Markdown code fence; json_schema structured output is model-specific | |
reasoning_content (From Deepseek API) | Text reasoning messages. The single response field; on an assistant message replayed in messages, reasoning is accepted as an alias for it, and it is dropped on models that only accept their own thinking | |
annotations (URL citations) | URL citations from a web search or another server tool (non-streaming only) | |
| Usage tracking | ||
| Input text tokens | Billing unit | |
| Output tokens | Billing unit | |
| Reasoning tokens | Converse-served models do not split them out: completion_tokens_details is not populated and reasoning tokens are billed inside completion_tokens. Mantle-native models report whatever split their upstream API returns | |
| Other | ||
| Service tiers | Mapped to Bedrock service tiers and latency options. fast is accepted as the alias of priority. The response reports the tier that served the request, which Amazon Bedrock names, so it can differ from the one asked for and reflects a tier the deployment configured in the request's place — see service tier configuration. When nothing names a tier, the requested one is reported after alias mapping, and any tier without a Bedrock equivalent as default. When streaming, the tier Amazon Bedrock reports only arrives with its last event: the content chunks carry the tier the call was sent on, and the trailing usage chunk the served one | |
metadata | Echoed in the response, updatable on stored completions, and usable to filter the Bedrock invocation log. Also forwarded to Bedrock requestMetadata, whose limits apply: max 16 pairs, values ≤256 characters, restricted character set. The gateway adds two stdapi-ai.* pairs of its own (three when an end user is identified), so send at most 14 (13) pairs | |
store | Persists the completion in Amazon Bedrock session storage (non-streaming) | |
| List / update stored completions | List with model/metadata filters; metadata update | |
safety_identifier / user | Does not affect generation; identifies the end user in the request log and in per-user cost attribution | |
| Bedrock Guardrails | Content safety policies — not applied to Mantle-served requests | |
moderation | Applies an Amazon Bedrock guardrail; results in the response moderation field, or, when streaming, in a dedicated moderation chunk with empty choices. Amazon Bedrock only reports the guardrail trace with its last event, so that chunk is sent last, after the usage one, and carries both directions at once — rejected (400) on Mantle-served models |
Legend:
- Supported — Fully compatible with OpenAI API
- Model-Dependent — Behavior depends on the model or backend; check the Notes column
- Partial — Supported with limitations
- Unsupported — Not available in this implementation
- Extra Feature — Enhanced capability beyond OpenAI API
Models¶
All models supported by the Amazon Bedrock Converse and Converse Stream API are supported, plus every model served by Bedrock Mantle when enabled — including OpenAI GPT-5.x, xAI Grok, and Google Gemma 4. Requests to Mantle models are passed through natively or converted automatically depending on the model's upstream API support — see Bedrock Mantle below.
Bedrock Mantle¶
Mantle-served requests follow one of three paths, each with its own parameter fidelity:
| Serving path | Models | Parameter behavior |
|---|---|---|
| Passthrough | Chat-native models (OpenAI GPT-6, xAI Grok, OpenAI gpt-oss, Google Gemma 4, other open-weight models) | All schema-accepted parameters are forwarded; the upstream API may reject unsupported ones per model with a clean 400 (the upstream error code and parameter are propagated) |
| Converted to Responses | OpenAI GPT-5.x frontier models; unknown models | Dropped silently: stop, seed, frequency_penalty, presence_penalty, logit_bias, top_logprobs, audio, modalities, message name, input_audio content parts, legacy functions/function_call. Preserved: metadata, safety_identifier. max_tokens/max_completion_tokens below 16 is raised to 16, the Responses API's minimum (a budget of 1, sent as a cheap model probe by some clients, would otherwise be rejected with 400); where the classic endpoint also serves the model (the GPT-5.6 family, by default), clearing AWS_BEDROCK_MANTLE_PREFERRED_MODELS moves it there, where a smaller budget is honored as sent. n > 1 rejected with 400. store is handled by stdapi.ai only, never forwarded upstream |
| Converted to Messages | Mantle-only Anthropic Claude models | Same drops and n > 1 rejection as the Responses conversion, except stop which is forwarded as stop_sequences; plus: temperature clamped to ≤ 1.0; max_tokens defaults to 4096 when unset; reasoning_effort mapped to Anthropic effort levels, or to a thinking budget on Claude 3.7 - 4.5 (thinking_budget and enable_thinking honoured; see Reasoning Control for small output limits); response_format json_object/json_schema not available; metadata, prompt_cache_key, and prompt_cache_retention dropped; service_tier forwarded only when auto; thinking text returned as reasoning_content |
Project attribution (OpenAI-Project)
Mantle requests can be attributed to a Bedrock Project for cost tracking and observability with the OpenAI-Project: <project-id> header (a bare project ID such as proj_abc123, not an ARN). It is honored per-request only when AWS_BEDROCK_ALLOW_MANTLE_PROJECT_OVERRIDE is true; otherwise the server default (AWS_BEDROCK_MANTLE_PROJECT) applies. This applies only to models served by the Bedrock Mantle endpoint — classic bedrock-runtime models ignore the header.
Model Name Aliases¶
This API supports dynamic model name aliases matching official provider APIs. Models like OpenAI and Anthropic provide dynamic aliases in their official APIs—this gateway supports the same model names, automatically resolving them to Amazon Bedrock model identifiers.
Examples (OpenAI GPT OSS models supported by Bedrock):
gpt-oss-20b→openai.gpt-oss-20b-1:0
Working with the Chat Completions API¶
Prompt Caching¶
Reduce costs and improve response times by caching frequently-used prompt components across multiple requests. This feature is particularly effective for applications with consistent system prompts, tool definitions, or conversation contexts.
Supported Models:
- Anthropic Claude: Full support for system, messages, and tools caching
- Amazon Nova: Support for system and messages caching
- Moonshot Kimi K3 and OpenAI GPT-6: Automatic — a repeated prompt prefix is cached with no parameter at all, and the cached tokens are reported in
usage.prompt_tokens_details.cached_tokens. Cache breakpoints are accepted and ignored.prompt_cache_keyis ignored too, except on GPT-6 served by Bedrock Mantle (its default where a configured Region offers it), where it is forwarded to the model as a cache routing key
Documentation
See Amazon Bedrock Prompt Caching - Supported Models for the complete list of models supporting prompt caching.
Cache Creation Costs
Cache creation incurs a higher cost than regular token processing. Only use prompt caching when you expect a high cache hit ratio across multiple requests with similar prompts.
How to Use:
Set the prompt_cache_key parameter to enable caching:
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-fable-5",
"prompt_cache_key": "default",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant with extensive knowledge..."
},
{"role": "user", "content": "What is 2 + 2?"}
]
}'
Granular Cache Control:
Enable caching for specific prompt sections using dot-separated values:
"system"- Cache system messages only"messages"- Cache conversation history"tools"- Cache tool/function definitions (Anthropic Claude only)"system.messages"- Cache both system and messages"system.tools"- Cache system and tools"messages.tools"- Cache messages and tools"system.messages.tools"- Cache all components- Any other non-empty value - Cache all components
Custom Cache Keys Not Supported
Custom cache hash keys are not supported. The parameter is used only to control which sections are cached, not as a cache identifier.
{
"model": "anthropic.claude-fable-5",
"prompt_cache_key": "system.tools",
"messages": [...],
"tools": [...]
}
Benefits:
- Cost Reduction: Cached tokens are billed at a lower rate than regular input tokens
- Lower Latency: Cached prompts eliminate reprocessing time
- Automatic Management: The API handles cache invalidation and updates
Cache Retention (TTL):
Control how long cached prompts persist using the prompt_cache_retention parameter:
Model Support
Cache retention configuration is only available on select models. See Amazon Bedrock Prompt Caching - Supported Models for details on which models support configurable TTL.
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-fable-5",
"prompt_cache_key": "default",
"prompt_cache_retention": "24h",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant..."
},
{"role": "user", "content": "What is 2 + 2?"}
]
}'
Available Retention Values:
"in_memory"- Short-term caching (mapped to 5 minutes on Amazon Bedrock)"24h"- Long-term caching (mapped to 1 hour on Amazon Bedrock)- Additional Amazon Bedrock values:
"1h","5m"(provider-specific)
OpenAI to Amazon Bedrock Mapping
OpenAI retention values are mapped to Amazon Bedrock equivalents for compatibility:
"in_memory"→ 5 minutes"24h"→ 1 hour
The OpenAI prompt_cache_options object is also accepted: its ttl ("30m") is mapped to a 1 hour Amazon Bedrock retention on Anthropic models (other models use the default 5 minute TTL) when prompt_cache_retention is not set.
Explicit Cache Breakpoints:
Instead of relying on the prompt_cache_key section heuristics, mark the exact cache boundaries with prompt_cache_breakpoint on any content part (text, image_url, input_audio, file, refusal). Each marked part is followed by an Amazon Bedrock cachePoint, so the prompt prefix ending with that part is cached:
{
"model": "anthropic.claude-fable-5",
"prompt_cache_options": {"mode": "explicit"},
"messages": [
{
"role": "system",
"content": [
{
"type": "text",
"text": "Long reusable instructions...",
"prompt_cache_breakpoint": {"mode": "explicit"}
}
]
},
{"role": "user", "content": "What is 2 + 2?"}
]
}
"mode": "explicit"caches only the marked parts: theprompt_cache_keyheuristics are disabled for that request."mode": "implicit"(default) keeps theprompt_cache_keyheuristics and honors the marked parts.- At most 4 cache points are sent per request (Amazon Bedrock limit); the oldest ones are dropped when more are requested.
- Breakpoints on models without prompt caching support are accepted and ignored, as are breakpoints on tool result messages — those never become a cache point, whatever the model — and a cache point that would directly follow a reasoning block, which Amazon Bedrock refuses.
Usage Tracking:
Cached token usage is reported in the response:
{
"usage": {
"prompt_tokens": 1500,
"completion_tokens": 100,
"total_tokens": 1600,
"prompt_tokens_details": {
"cached_tokens": 1200,
"cache_write_tokens": 300
}
}
}
In this example, 1,200 tokens were retrieved from cache and the remaining 300 tokens were processed and written to the cache. cache_write_tokens (an extra field beyond the OpenAI API) reports the tokens written to the cache when the model reports them, on both non-streaming responses and the trailing stream_options.include_usage chunk; it is omitted when no cache write occurred.
System Prompt¶
System prompts define the AI assistant's behavior, personality, and instructions (e.g., "You are a helpful assistant"). Most models support system prompts.
Unsupported Models
Some models don't support system prompts (mistral.mistral-7b-instruct-v0:2, mistral.mixtral-8x7b-instruct-v0:1). By default, stdapi.ai silently drops system messages for these models, allowing cross-model compatibility. To receive errors instead, configure DROP_UNSUPPORTED_SYSTEM_PROMPT=false.
S3 Image Support¶
Access images directly from your S3 buckets without generating pre-signed URLs or downloading files locally.
Supported Formats:
- Images: JPEG, PNG, GIF, WebP
How to Use:
Simply reference your S3 images using the s3:// URI scheme in image_url fields:
{
"model": "anthropic.claude-fable-5",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image"},
{
"type": "image_url",
"image_url": {"url": "s3://my-bucket/images/photo.jpg"}
}
]
}
]
}
IAM Permissions Required
Your API service must have IAM permissions to read from the specified S3 buckets. S3 objects must be in the same AWS region as the executed model or accessible via your IAM role. Standard S3 data transfer and request costs apply.
Benefits:
- No pre-signed URLs - Direct S3 access without generating temporary URLs
- Security - Images stay in your AWS account with IAM-controlled access
- Performance - Optimized data transfer within AWS infrastructure
- Large images - No size limitations of data URIs or base64 encoding
Files API References (file-id:)¶
The string-overloaded image_url.url, file.file_data, and input_audio.data fields also accept the project-local file-id: URI scheme to reference files previously uploaded via the Files API:
{
"model": "anthropic.claude-fable-5",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image"},
{"type": "image_url", "image_url": {"url": "file-id:file-0190c51c7de7455d9b8c2efe27dfbf67"}}
]
}]
}
Two equivalent ways to reference an uploaded file
The OpenAI-native typed path {"type": "file", "file": {"file_id": "file-…"}} is unchanged and still preferred when a typed object is acceptable. The file-id: URI is the equivalent for the string-overloaded image_url.url / file.file_data / input_audio.data fields, where today you would otherwise pass an s3://, https://, or data: URI. See Referencing Uploaded Files via the file-id: URI Scheme.
Server Tools¶
Amazon Bedrock system tools are built-in capabilities that foundation models can use directly, without requiring you to implement backend integrations.
How to Use:
Add system tools to your tools array as normal. System tools don't require parameter definitions—just specify the tool name and the model will handle the rest.
Amazon Nova Tools¶
| Tool | function.name | Amazon Nova 2 | Amazon Nova Premier (legacy) | API Support |
|---|---|---|---|---|
| Web Grounding | nova_grounding | |||
| Code Interpreter | nova_code_interpreter |
Code Interpreter Not Compatible
nova_code_interpreter cannot be used via this API. The code execution result cannot be surfaced in the OpenAI Chat Completions response format.
Web Grounding¶
Amazon Nova Web Grounding enables models to search the web for current information, helping answer questions requiring real-time data like news, weather, product availability, or recent events. The model automatically determines when to use web grounding based on the user's query.
Learn More
Usage: set web_search_options, the standard OpenAI parameter for a grounded answer. An empty object is enough.
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "amazon.nova-2-lite-v1:0",
"messages": [
{
"role": "user",
"content": "What are the current AWS Regions and their locations?"
}
],
"web_search_options": {}
}'
Naming nova_grounding as a function tool does the same thing, and is what a client that only speaks tools can send. Sending both declares the search once.
{"tools": [{"type": "function", "function": {"name": "nova_grounding"}}]}
Response Format:
When using web grounding, the API response includes annotations with URL citations in non-streaming mode:
{
"choices": [{
"message": {
"role": "assistant",
"content": "The AWS Regions include...",
"annotations": [
{
"type": "url_citation",
"url_citation": {
"url": "https://aws.amazon.com/about-aws/global-infrastructure/",
"title": "AWS Global Infrastructure"
}
}
]
}
}]
}
Streaming Mode
URL citation annotations are only available in non-streaming responses.
Limitations:
- No streaming citations: URL citation
annotationsare not emitted in streaming responses. web_search_optionssub-fields:search_context_sizeis accepted and ignored — the search runs with the amount the model applies.user_locationis rejected with400: the search cannot be restricted to a location, and running it worldwide would answer something other than what was asked.- Models without a web search:
web_search_optionsis rejected with400on every model that runs no web search of its own, including Mantle-served ones. Use one of the models above, or theweb_searchtool of the Responses API.
Region Compatibility
Web Grounding is only available in US Amazon Bedrock regions. To ensure all requests are routed to a US region, restrict the model using AWS_BEDROCK_MODEL_REGION_RESTRICT:
export AWS_BEDROCK_MODEL_REGION_RESTRICT='{"amazon.nova-": ["us-east-1"]}'
Anthropic Claude Server Tools¶
Anthropic Claude models support server-side tools (bash, text editor, memory) that are executed by the model provider. Declare them using the standard OpenAI function tool format: set type to "function" and function.name to the tool name.
Supported Tools by Model:
| Tool | function.name | Claude 3.5 Sonnet v2 | Claude 3.7+ |
|---|---|---|---|
| Bash | bash | ||
| Text Editor | str_replace_based_edit_tool | ||
| Computer | computer | ||
| Memory | memory |
Computer Use Not Supported
The computer use workflow requires screenshots to be returned as images inside tool results. The OpenAI Chat Completions API does not support image content in role: "tool" messages, so the complete agent loop cannot be implemented. computer is not usable via this route.
Usage:
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-sonnet-5",
"messages": [
{"role": "user", "content": "Run a Python script that prints hello world."}
],
"tools": [
{"type": "function", "function": {"name": "bash"}},
{"type": "function", "function": {"name": "str_replace_based_edit_tool"}}
]
}'
Tool Parameters:
Some Claude server tools accept additional configuration. Pass tool-specific parameters inside function.parameters:
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-sonnet-5",
"messages": [
{"role": "user", "content": "Edit the file hello.py to print hello world."}
],
"tools": [
{
"type": "function",
"function": {
"name": "str_replace_based_edit_tool",
"parameters": {"type": "object", "max_characters": 5000}
}
}
]
}'
Beta Headers
Claude server tools require specific anthropic-beta flags, which are automatically injected — no manual header needed:
bash,str_replace_based_edit_tool→computer-use-2024-10-22(Claude 3.5) orcomputer-use-2025-01-24(Claude 3.7+)memory→context-management-2025-06-27(Claude 3.7+)
Reasoning Control¶
This API supports several approaches to control Amazon Bedrock reasoning behavior. Reasoning enables foundation models to break down complex tasks into smaller steps ("chain of thought"), improving accuracy for multi-step analysis, math problems, and complex reasoning tasks. All approaches work with all Amazon Bedrock models that support reasoning capabilities.
Model Support for Configurable Reasoning
Not all reasoning-capable models support configurable reasoning control. Support varies by model:
- Anthropic Claude 3.7 - 4.5: Both
reasoning_effortandthinking_budgetparameters supported (token budget-based reasoning) - Anthropic Claude Sonnet 4.6 / Opus 4.6 and later (including Fable and Mythos):
reasoning_effortparameter only (adaptive reasoning) - Amazon Nova 2 models:
reasoning_effortparameter only - DeepSeek V3 models:
reasoning_effortparameter only - OpenAI GPT-5.x and GPT-6:
reasoning_effortparameter only - OpenAI gpt-oss:
reasoning_effortparameter only, on alow/medium/highscale - Moonshot Kimi K3:
reasoning_effortparameter only (the Moonshotthinkingtoggle below also turns reasoning off)
Models listed as effort-only still accept a token budget: it turns reasoning on, and the depth comes from their own effort scale.
A small output limit turns reasoning off on Claude 3.7 - 4.5
These models spend a thinking budget of at least 1,024 tokens out of the output limit. A reasoning_effort sent with a max_completion_tokens (or max_tokens) of 1,024 or less leaves no room for it, so the request is served without reasoning and a warning is logged. Raise the limit above 1,024 to get reasoning back.
OpenAI and DeepSeek API-Compatible Reasoning Parameters¶
Use the reasoning_effort parameter with predefined effort levels. This format is shared by the OpenAI and DeepSeek Chat Completions APIs and works with all Amazon Bedrock models supporting reasoning.
Available Levels:
none- Disable reasoning (sent asbetween_tools, their lowest setting, to Claude Sonnet 5.5 and later; accepted but not honored on Claude Opus 5.5 and later, Fable and Mythos, which always reason: the model's default adaptive mode is used, the response still carries reasoning content and its output tokens are still billed; the same applies to OpenAI GPT-6 Astra and GPT-6.1 and later, which reason at their default level, and to gpt-oss, except on Bedrock Mantle, where gpt-oss receivesnoneas sent)minimal- Quick responses with minimal reasoning (sent aslowto the models without a minimal level: OpenAI GPT-5.x and GPT-6 wherever they are served, and gpt-oss and Moonshot Kimi outside Bedrock Mantle)low- Light reasoning for straightforward tasksmedium- Balanced reasoning for most use caseshigh- Deep reasoning for complex problemsxhigh- Maximum reasoning for complex problemsmax- Its own (higher) effort tier on Claude Sonnet/Opus 4.6 and later, Fable and Mythos, and on OpenAI GPT-5.x, GPT-6 and Moonshot Kimi K3, which receive it unchanged; collapsed onto the model's top reasoning tier on the fixed-scale models (Claude 3.7 - 4.5, Amazon Nova 2, DeepSeek, Kimi K2, gpt-oss). Claude 4.6 mapsxhighdown tohigh. OpenAI GPT-5.x and GPT-6 get theminimalandnonehandling above wherever they are served; other models served through Bedrock Mantle receive the level as sent
What You Get:
reasoning_contentfield (DeepSeek API-compatible): models include their thinking process in the response. The field name is an operator setting (CHAT_COMPLETIONS_REASONING_FIELD):reasoning_contentby default,reasoningfor clients written against OpenRouter or vLLM, ornoneto keep responses strictly OpenAI-shaped.- Streaming support:
choices[].delta.reasoning_contentchunks in real time - No text from recent Claude models: Claude Opus 4.7 and later, Sonnet 5, Fable and Mythos reason without returning the text unless asked for a summary, which this API cannot request, so their reasoning field stays empty while its tokens are billed. Read their reasoning through Messages
thinking.displayor Responsesreasoning.summary
Example:
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-fable-5",
"reasoning_effort": "high",
"messages": [{"role": "user", "content": "Solve this complex problem..."}]
}'
Compatibility
This format is accepted for all reasoning-capable models. Models that don't support this parameter will ignore it.
Qwen API-Compatible Reasoning Parameters¶
Use explicit enable_thinking & thinking_budget parameters for fine-grained control over thinking mode. This is a Qwen API-compatible format that works with all Amazon Bedrock models supporting reasoning.
Parameters:
enable_thinking(boolean): Enable or disable thinking mode- Default: Model-specific (usually
false)
- Default: Model-specific (usually
thinking_budget(integer): Maximum thinking process length in tokens- Only effective when
enable_thinkingistrue - Passed to the model as
budget_tokens - Default: Model's maximum chain-of-thought length
- Only effective when
Example:
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-sonnet-5",
"enable_thinking": true,
"thinking_budget": 2000,
"messages": [{"role": "user", "content": "Solve this complex problem..."}]
}'
Compatibility
This format is accepted for all reasoning-capable models. Models that don't support these parameters will ignore them.
Moonshot API-Compatible Thinking Control¶
The thinking parameter provides a Moonshot API-compatible format for controlling thinking/reasoning on models that support it.
Documentation
See Moonshot Kimi API for more information.
Parameters:
thinking={"type": "enabled"}— Enable thinking modethinking={"type": "disabled"}— Disable thinking mode
This format is accepted for all reasoning-capable models. Models that don't support this parameter will ignore it.
OpenRouter API-Compatible Reasoning Object¶
The reasoning object groups the same controls into a single field, and include_reasoning toggles whether the reasoning text is returned.
Documentation
See OpenRouter Reasoning Tokens for more information.
Parameters:
reasoning.effort(string): Effort level, exactly asreasoning_effortreasoning.max_tokens(integer): Reasoning token budget, exactly asthinking_budget(impliesenabled)reasoning.enabled(boolean): Enable reasoning, exactly asenable_thinkingreasoning.exclude(boolean): Omit the reasoning text from the responseinclude_reasoning(boolean):falseis equivalent toreasoning: {"exclude": true}
Example:
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-sonnet-5",
"reasoning": {"effort": "high", "exclude": true},
"messages": [{"role": "user", "content": "Solve this complex problem..."}]
}'
Conflicting Values
reasoning.effort and reasoning.max_tokens are mutually exclusive, and a sub-field disagreeing with its flat equivalent (for example reasoning: {"effort": "high"} with reasoning_effort: "low") is rejected with 400.
Excluded Reasoning
Excluding the reasoning text does not disable reasoning: the model still thinks, and the reasoning tokens are still counted in usage and billed. Use reasoning_effort: "none" or reasoning: {"enabled": false} to turn reasoning off.
Replaying Reasoning in a Multi-Turn Conversation¶
Appending the assistant message you just received to messages — the standard multi-turn idiom, and what the DeepSeek API asks for after a tool call — is always accepted, reasoning_content included.
Anthropic Claude models only continue from a thinking passage they can recognise as their own, which a replayed text field is not, so their reasoning is left out of that turn instead. The message content, tool calls and refusal are sent unchanged and the conversation continues normally; only the earlier chain of thought is no longer visible to the model. Every other model family receives the replayed reasoning as-is.
What this does not affect, measured on Claude Haiku 4.5:
- Reasoning on the new turn. The model still thinks, and still returns
reasoning_content— the reasoning setting for the turn being generated is independent of the history. It re-derives rather than continuing the earlier chain. - Tool-call continuations. A turn that carried a tool call is answered correctly with the earlier reasoning left out.
- Prompt caching. Cache hits are unaffected, including when the cache point sits inside the conversation immediately after the turn whose reasoning is left out — measured across three turns, each one reading the previous turn's cache in full and extending it. The omission is the same on every turn, so the cached prefix stays identical and keeps growing.
Keeping the reasoning in context on Claude
Use the Responses API, which carries the thinking passage in a form Claude accepts. Request include: ["reasoning.encrypted_content"] and echo the reasoning items back, and the model continues from its own earlier reasoning.
Provider-Specific Parameters¶
Unlock advanced model capabilities by passing provider-specific parameters directly in your requests. These parameters are forwarded to Amazon Bedrock and allow you to access features unique to each foundation model provider.
Documentation
See Bedrock Model Parameters for the complete list of available parameters per model.
How It Works:
Add provider-specific fields at the top level of your request body alongside standard OpenAI parameters. The API automatically forwards these to the appropriate model provider via Amazon Bedrock.
Examples:
Top K Sampling:
{
"model": "anthropic.claude-fable-5",
"messages": [{"role": "user", "content": "Write a poem"}],
"top_k": 50,
"temperature": 0.7
}
Configuration Options:
Option 1: Per-Request
Add provider-specific parameters directly in your request body (as shown in examples above).
Option 2: Server-Wide Defaults
Configure default parameters for specific models via the DEFAULT_MODEL_PARAMS environment variable:
export DEFAULT_MODEL_PARAMS='{
"anthropic.claude-sonnet-5": {
"anthropic_beta": ["context-management-2025-06-27"]
}
}'
Parameter Priority
Per-request parameters override server-wide defaults.
Behavior:
- Compatible parameters: Forwarded to the model and applied
- Unsupported parameters: Return HTTP 400 with an error message
- Reserved names:
model_idandadditional_request_fieldscollide with the gateway's own request-building parameters and are rejected with a400 invalid_request_errornaming the key, instead of being forwarded
Anthropic Claude Features¶
Enable cutting-edge Claude capabilities through Anthropic beta feature flags.
Beta Feature Flags¶
Enable experimental Claude features like interleaved thinking by adding the anthropic_beta array to your request (extended thinking itself is controlled through the reasoning parameters, not a beta flag):
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-sonnet-5",
"messages": [{"role":"user","content":"Summarize the news headline."}],
"anthropic_beta": ["interleaved-thinking-2025-05-14"]
}'
Server-Wide Configuration
You can also configure beta flags server-wide using the DEFAULT_MODEL_PARAMS environment variable (see Provider-Specific Parameters).
Unsupported Beta Flags
Unsupported flags that would change output return HTTP 400 errors.
Documentation
See Using Claude on Amazon Bedrock for more details on Claude-specific parameters.
Stored Chat Completions¶
Set store: true to persist a chat completion in Amazon Bedrock session storage — same mechanism, region, and KMS setting as stored responses. The returned id then works with the full stored-completion surface:
GET /v1/chat/completions— list stored completions, sorted by creation time (order,after,limit), filterable bymodeland by metadata pairs. Metadata filters accept either onemetadata[key]=valueparameter per key (what the OpenAI SDKs send) or a singlemetadata={"key": "value"}JSON object of string values, for clients that can only send a whole object in one query parameter. A baremetadatain any other shape is rejected with400naming both accepted forms.GET /v1/chat/completions/{completion_id}— retrieve the stored completion.POST /v1/chat/completions/{completion_id}— replace itsmetadata(nullclears it).GET /v1/chat/completions/{completion_id}/messages— list its input messages.DELETE /v1/chat/completions/{completion_id}— delete it and its backing session.
store defaults to false on this implementation and is ignored with stream=true or when the server lacks the session storage IAM permissions (a warning is recorded in the request log). Listings scan a capped number of sessions (1,000) in the primary Bedrock region; accounts beyond the cap may see incomplete listings.
Model-Specific Features¶
TwelveLabs Pegasus¶
twelvelabs.pegasus-1-2-v1:0 is a video-understanding model. Because Pegasus accepts exactly one video and one text prompt per call, this API adapts the conversation automatically:
- The latest video found anywhere in the conversation (any role, any position) is forwarded as the video input.
- The latest contiguous run of user text (back to the previous assistant or tool turn) is concatenated and forwarded as the text prompt.
temperatureandmax_tokensare forwarded.response_format: json_schemais forwarded as Pegasus's structured output.
Silently ignored (no error): system prompts, tools, top_p, stop sequences, and prompt caching.
Upstream format limitation: The OpenAI Chat Completions API has no video_url content part type. To stay fully compatible with standard OpenAI clients, pass the video as an image_url content part — the server detects the video MIME type automatically and routes it to Pegasus correctly.
Video input formats: data:video/mp4;base64,…, https://…, s3://bucket/key, or file-id:…. Videos above 18.75 MB are automatically uploaded to S3.
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "twelvelabs.pegasus-1-2-v1:0",
"messages": [
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {"url": "s3://my-bucket/video.mp4"}
},
{"type": "text", "text": "Describe what happens in this video."}
]
}
]
}'
Limits and behaviour to know¶
What is rejected with a 400. logprobs when enabled, web_search_options on a model that runs no web search (and its user_location on every model), translation_options, custom (free-form) tools, tool_choice: {"type": "allowed_tools"}, a whitespace-only stop sequence (an Amazon Bedrock limitation), audio output combined with stream: true, and n > 1 combined with stream: true. Each error names the parameter it refuses.
What is accepted and ignored. prediction and verbosity are latency and length hints the Converse API has no equivalent for, the message name participant field has no Converse Message author field to land in, web_search_options.search_context_size only tunes a search that still runs, and stream_options.include_obfuscation never pads the stream. prediction, verbosity, name and stream_options.include_obfuscation are forwarded verbatim on Mantle passthrough models, where the upstream API decides what to do with them.
temperature above 1.0 is served at 1.0. The documented 0–2 range is accepted in full: a higher value is applied at the maximum Amazon Bedrock accepts rather than refused, so 1.5 and 2.0 sample as 1.0 does. Mantle passthrough models receive the value as sent.
What changes on a Mantle-served model. An Amazon Bedrock guardrail is not applied to a Mantle-served request, the moderation and web_search_options parameters are rejected with 400 (ask for the search with the web_search tool of the Responses API instead), and audio synthesis from text output is not performed. Where the model is not chat-native, the request is converted to the Responses or Messages API and a documented set of parameters is dropped or clamped — the three serving paths and their exact parameter fidelity are tabled under Bedrock Mantle.
Reasoning tokens are not split out on Converse-served models. completion_tokens_details is not populated and reasoning tokens are billed inside completion_tokens; Mantle-native models report whatever split their upstream API returns.
Request headers¶
This endpoint supports standard Bedrock headers for enhanced control over your requests. All headers are optional and can be combined as needed.
Content Safety (Guardrails)¶
| Header | Purpose | Valid Values |
|---|---|---|
X-Amzn-Bedrock-GuardrailIdentifier | Guardrail ID for content filtering | Your guardrail identifier |
X-Amzn-Bedrock-GuardrailVersion | Guardrail version | Version number (e.g., 1) |
X-Amzn-Bedrock-Trace | Guardrail trace level | disabled, enabled, enabled_full |
X-Amzn-Bedrock-GuardrailStreamProcessingMode | Guardrail assessment timing on a streamed response | sync, async (streaming requests only) |
Performance Optimization¶
| Header | Purpose | Valid Values |
|---|---|---|
X-Amzn-Bedrock-Service-Tier | Service tier selection | default, flex, priority, reserved |
X-Amzn-Bedrock-PerformanceConfig-Latency | Latency optimization | standard, optimized |
Example with all headers:
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-H "X-Amzn-Bedrock-GuardrailIdentifier: your-guardrail-id" \
-H "X-Amzn-Bedrock-GuardrailVersion: 1" \
-H "X-Amzn-Bedrock-Trace: enabled" \
-H "X-Amzn-Bedrock-Service-Tier: priority" \
-H "X-Amzn-Bedrock-PerformanceConfig-Latency: optimized" \
-d '{
"model": "anthropic.claude-fable-5",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Detailed Documentation
For complete information about these headers, configuration options, and use cases, see:
Try it¶
Streaming response:
curl -N -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "amazon.nova-micro-v1:0",
"stream": true,
"messages": [{"role": "user", "content": "Write a haiku about the sea."}]
}'
Multi-modal with image:
{
"model": "amazon.nova-micro-v1:0",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image"},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}
]
}
]
}
With reasoning:
curl -X POST "$BASE/v1/chat/completions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-sonnet-5",
"reasoning_effort": "low",
"messages": [{"role": "user", "content": "Solve 12*13"}]
}'
Response with reasoning:
{
"choices": [{
"message": {
"role": "assistant",
"reasoning_content": "12 × 10 = 120, plus 12 × 3 = 36 → 156",
"content": "156"
}
}]
}
Next steps¶
Next: Responses API · Models API · Files API · Moderations API