Chat API (Ollama Compatible)¶
Generate conversational AI responses with Amazon Bedrock models through the Ollama /api/chat interface. Served under /ollama by default; the examples below use $BASE, which includes that prefix.
At a glance¶
- Drop-in Ollama compatibility — Follows the Ollama
/api/chatrequest and response shape, including its newline-delimited JSON streaming transport, so an existing Ollama client works by changing the base URL. - Tool calling and thinking — Function tools and thinking (
think) work the same way they do against a local Ollama server. - Structured output —
formataccepts"json"or a full JSON Schema, and the answer is constrained to it. - Private AWS backend — Served entirely by Amazon Bedrock models in your own AWS account — no traffic to third-party endpoints.
- Image input beyond base64 —
messages[].imagesalso takes a URL, a data URI or ans3://URI on models that read images. - Differs from the Ollama API: models are never resident, so
keep_aliveholds nothing loaded andload_durationis never reported;logprobsis refused with400; runner options such asnum_ctxare accepted and ignored — see Limits and behaviour to know.
Base URL and route prefix
By default, all Ollama-compatible routes are prefixed with /ollama. This means the Chat API is available at /ollama/api/chat instead of /api/chat. You can customize this prefix using the OLLAMA_ROUTES_PREFIX configuration variable documented in HTTP Server and MCP.
The curl examples on this page use a $BASE variable that must include this prefix — set it to your scheme and host followed by OLLAMA_ROUTES_PREFIX:
export BASE="https://your-host/ollama" # <scheme>://<host> + OLLAMA_ROUTES_PREFIX
curl -X POST "$BASE/api/chat" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "amazon.nova-micro-v1:0",
"messages": [{"role": "user", "content": "Say hello world"}],
"stream": false
}'
Endpoints¶
| Endpoint | Method | What It Does | Powered By | MCP Tool |
|---|---|---|---|---|
/api/chat | POST | Conversational AI, following the Ollama Chat API | Amazon Bedrock chat models | ollama_chat |
Model Names¶
Send the model names GET /api/tags publishes — those are the canonical identifiers this server resolves directly. A trailing :latest is accepted and stripped as a fallback when the exact name is not found. Short aliases accepted on this server's other APIs also work here even though /api/tags does not list them. A name learned from ollama.com (for example llama3.2:3b) is not available through this server and answers 404. Every response echoes the model name exactly as the request spelled it, in model.
Find compatible models: Call /search_models with route=ollama_chat to discover model IDs that support this route, or call GET /api/tags for the Ollama-shaped listing.
Feature compatibility¶
| Feature | Status | Notes |
|---|---|---|
| Input | ||
messages (text) | Full support | |
messages[].images | Multimodal models only; base64, a URL, a data URI or an s3:// URI | |
messages[].thinking | Replayed as the assistant turn's reasoning text | |
messages[].tool_calls | Replayed tool calls; correlated to results as described below | |
tools | Function tools; support depends on the model | |
format | "json" or a JSON Schema object; see Structured Output | |
options | temperature, top_p, top_k, seed, stop, num_predict (max output tokens), presence_penalty and frequency_penalty are forwarded; runner options (num_ctx, num_gpu, num_thread, num_batch, main_gpu, use_mmap, min_p, and any other unknown key) are accepted and ignored | |
stream | Newline-delimited JSON; defaults to true — see Streaming | |
think | Boolean, or low/medium/high/max; see Thinking | |
keep_alive | Ignored for residency — models are never resident — but still tells a request with no message apart as a load or an unload, below | |
logprobs / top_logprobs | Rejected with 400 | |
| Output | ||
message.content | Full support | |
message.thinking | Follows CHAT_COMPLETIONS_REASONING_FIELD; omitted when the operator sets that to none | |
message.tool_calls | Streamed whole in one event, never as partial argument fragments — see Tool Calling | |
done_reason | stop or length on a generated answer; load or unload on a message-less request — see Loading and Unloading | |
| Usage tracking | ||
prompt_eval_count, eval_count | Real token counts | |
total_duration | Real wall-clock time | |
load_duration | Never reported — there is no model-loading phase to measure | |
prompt_eval_duration, eval_duration | Reported only when streaming, measured from the stream itself; omitted on a non-streaming response |
Legend:
- Supported — Fully compatible with the Ollama API
- Available on Select Models — Check your model's capabilities
- Partial — Supported with limitations
- Unsupported — Not available in this implementation
Streaming¶
By default (stream unset, or true), the response body is newline-delimited JSON (application/x-ndjson): one bare JSON object per line, with no data: prefix and no [DONE] sentinel. The stream is terminated by an object carrying "done": true and the response metrics.
curl -N -X POST "$BASE/api/chat" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "amazon.nova-micro-v1:0",
"messages": [{"role": "user", "content": "Write a haiku about the sea."}]
}'
{"model":"amazon.nova-micro-v1:0","created_at":"2026-08-27T12:00:00.000000+00:00","message":{"role":"assistant","content":"Waves"},"done":false}
{"model":"amazon.nova-micro-v1:0","created_at":"2026-08-27T12:00:00.100000+00:00","message":{"role":"assistant","content":" crash on silent shores"},"done":false}
{"model":"amazon.nova-micro-v1:0","created_at":"2026-08-27T12:00:00.200000+00:00","message":{"role":"assistant","content":""},"done":true,"done_reason":"stop","total_duration":812345678,"prompt_eval_count":9,"prompt_eval_duration":123456789,"eval_count":11,"eval_duration":688888889}
A Streamed Failure Is Not an HTTP Status
Once the stream has begun, a failure can no longer be reported as an HTTP status: the response headers were already sent with a 200. Instead, the stream ends with a final {"error": "<message>"} line in place of the terminal done: true object. Check every line for an error key rather than relying on the status code alone.
Tool Calling¶
Ollama tool calls carry no identifier of their own. When replaying a conversation, messages[].tool_calls entries are matched to the following tool message that answers them by tool_call_id when the client sent one, then by tool_name, then in call order.
Streamed tool calls are emitted whole, in a single event once the model has finished requesting them, rather than as partial argument fragments — matching what Ollama itself does.
{"model":"amazon.nova-micro-v1:0","created_at":"2026-08-27T12:00:00.000000+00:00","message":{"role":"assistant","content":"","tool_calls":[{"function":{"name":"get_weather","arguments":{"city":"Paris"},"index":0}}]},"done":false}
Structured Output¶
Set format to "json" for unstructured JSON mode, or to a JSON Schema object for validated structured output:
curl -X POST "$BASE/api/chat" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-sonnet-5",
"messages": [{"role": "user", "content": "Extract the city and country."}],
"format": {
"type": "object",
"properties": {"city": {"type": "string"}, "country": {"type": "string"}},
"required": ["city", "country"]
},
"stream": false
}'
An object schema that does not set additionalProperties is closed automatically (additionalProperties: false), so a schema written for a local Ollama server works unchanged here.
A schema constrains the answer only on models whose backend accepts one; a model that does not answers 400 naming the parameter. "json" mode is more widely available. Filter with /search_models or try the request — the model is the authority.
Thinking¶
Set think to true, or to low, medium, high or max, to request the model's reasoning. The reasoning text comes back in message.thinking, both in the final response and on the streaming deltas that carry it.
curl -X POST "$BASE/api/chat" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic.claude-sonnet-5",
"think": "high",
"messages": [{"role": "user", "content": "Solve 12*13"}],
"stream": false
}'
{
"model": "anthropic.claude-sonnet-5",
"created_at": "2026-08-27T12:00:00.000000+00:00",
"message": {
"role": "assistant",
"thinking": "12 x 10 = 120, plus 12 x 3 = 36 -> 156",
"content": "156"
},
"done": true,
"done_reason": "stop",
"total_duration": 950123456,
"prompt_eval_count": 14,
"eval_count": 22
}
message.thinking follows the CHAT_COMPLETIONS_REASONING_FIELD server setting: when an operator sets it to none, no thinking text is emitted on this API either, whatever think was sent.
Images¶
messages[].images accepts base64-encoded image data, as Ollama does. This server additionally accepts a URL, a data URI or an s3:// URI in the same field, on models that support image input.
Loading and Unloading¶
A request with an empty messages array is upstream's way of making a model resident, and the same request with keep_alive set to 0 is how it is evicted — what a client's "load model" and "unload model" controls send. A hosted model is always resident, so both are answered without invoking anything:
{
"model": "amazon.nova-micro-v1:0",
"created_at": "2026-01-01T00:00:00+00:00",
"message": { "role": "assistant", "content": "" },
"done": true,
"done_reason": "load"
}
done_reason is unload when keep_alive is 0, load otherwise. The answer is a single JSON object whatever stream says, as upstream's is.
Limits and behaviour to know¶
logprobsandtop_logprobsare rejected with400— log probabilities are not available.keep_alivedoes not keep anything loaded: models are never resident. It is read only to tell a message-less request'sdone_reasonapart,loadfromunload.- Runner options inside
options(num_ctx,num_gpu,num_thread,num_batch,main_gpu,use_mmap,min_p, and any other key a local runner would use) are accepted and ignored. load_durationis never reported: there is no model-loading phase to measure, and a number there would be invented.prompt_eval_durationandeval_durationare reported only when streaming. All duration and count fields are optional in the Ollama API, so a client computing tokens-per-second from a non-streamed response has no duration to divide by.- A model name learned from ollama.com —
llama3.2:3b, for one — names nothing this server serves and answers404; send a nameGET /api/tagspublishes.
Request headers¶
| Header | Purpose | Notes |
|---|---|---|
Authorization | Gateway API key | Bearer <key>, required like every other route |
A local Ollama server needs no key; this one does, on every route.
Try it¶
curl -X POST "$BASE/api/chat" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "amazon.nova-micro-v1:0",
"messages": [{"role": "user", "content": "Say hello world"}],
"stream": false
}'
Example response:
{
"model": "amazon.nova-micro-v1:0",
"created_at": "2026-08-27T12:00:00.000000+00:00",
"message": {"role": "assistant", "content": "Hello, world!"},
"done": true,
"done_reason": "stop",
"total_duration": 812345678,
"prompt_eval_count": 6,
"eval_count": 5
}
Next steps¶
Next: Generate API · Embed API · Models API · Search Models API