Skip to content

Batch API

Run a large set of chat completion or embedding requests asynchronously, at a lower price than the synchronous API, through the OpenAI Batch API shape.

A batch is created from a file of requests, runs without a connection held open, and is read back from the result files it produces — exactly the OpenAI workflow, so the official OpenAI SDKs work by changing the base URL.

At a glance

  • Four endpoints — create, list, retrieve and cancel, served by Amazon Bedrock batch inference, see Endpoints.
  • The batch rate, roughly half the on-demand rate — usage is recorded once, when the batch ends, see Billing.
  • 100 to 50,000 requests per batch, in up to 200 MB of JSONL — a batch outside those bounds, or holding a line that cannot be run as it stands, is refused when it is created, see What the input file is checked for.
  • A 24-hour processing window, no connection held — completion_window is 24h as upstream, counted from created_at and reported as expires_at, see Listing order.
  • Requests and results in your own S3 bucket — the results are read back through the Files API as output_file_id and error_file_id, with no traffic to third-party endpoints.
  • No tools, json_schema, streaming or n above 1 — a batch carrying one of them is refused when it is created, see Feature compatibility.
  • A few creations at a time — the create call reads and translates the whole file, so a caller past that bound is answered 429 or 503 rather than queued, see How many batches may be created at once.
  • No DELETE /v1/batches/{batch_id} — deletion is served by the Anthropic Message Batches surface, for the batches created there, see Deleting a batch.
curl -X POST "$BASE/v1/batches" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "input_file_id": "file-06fvfg3lbdqarbad8kbo55g0sg5h3s4a",
    "endpoint": "/v1/chat/completions",
    "completion_window": "24h"
  }'

Endpoints

Endpoint Method What It Does MCP Tool
/v1/batches POST Create a batch from an uploaded file openai_batch
/v1/batches GET List batches, newest first openai_batch_list
/v1/batches/{batch_id} GET Retrieve a batch and its counters openai_batch_get
/v1/batches/{batch_id}/cancel POST Cancel a batch that is still running openai_batch_cancel

Feature compatibility

Feature Status Notes
Creation
input_file_id Must be uploaded with purpose="batch"
endpoint /v1/chat/completions and /v1/embeddings
completion_window 24h, as upstream
metadata Up to 16 key-value pairs, returned on every read
output_expires_after anchor: "created_at" and 1 hour to 30 days, counted from the moment the result files are written; omit it to keep them until deleted
Per-request body
messages, max_tokens, sampling Same parameters as Chat Completions
response_format json_object Full support
tools / tool_choice / functions Refused when the batch is created — tool use is not available in a batch
response_format json_schema Refused when the batch is created
stream A batch has nothing to stream to
n above 1 Send one request per completion instead
service_tier Accepted and ignored — a batch always runs on Amazon Bedrock batch inference, and the completions report no service_tier because the OpenAI schema publishes no value for it; the usage is billed at the batch rate
prompt_cache_key / prompt_cache_breakpoint Accepted and ignored — a batch reads and writes no prompt cache, and the request is answered without one
input, dimensions Same parameters as Embeddings, one input per request
encoding_format base64 Refused when the batch is created — batched vectors come back as numbers
Lifecycle
Retrieve / poll validating → in_progress → finalizing → completed
failed status and errors Reported for a batch that was accepted and then could not run at all, with one batch-level entry in errors.data[]; a problem in the input file is refused when the batch is created instead — see What the input file is checked for
Cancel cancelling then cancelled; requests already answered stay in output_file_id, and a batch that has ended is unchanged
List batches Newest first by created_at, with an after cursor — see Listing order
output_file_id / error_file_id Readable through the Files API
usage Token totals, reported once the batch ends
finalizing status Reported with finalizing_at while the results of a batch whose requests have run are being assembled; completed follows once they are readable

Legend:

  • Supported — Fully compatible with OpenAI API
  • Partial — Supported with limitations
  • Unsupported — Not available in this implementation

Content Guardrails and Batches

A request that a guardrail would apply to is refused rather than run unguarded. Send those requests without batching.

Prompt Caching and Batches

Batched requests neither read nor write a prompt cache, on any model. A request carrying a cache hint is still accepted and answered — the hint is dropped rather than the request — so a batch reports no cached tokens in usage.input_tokens_details. Nothing is lost by leaving the hint in: batched requests are already billed at the batch rate, and the cache discount was never available at that rate.

Models

Any chat or embedding model available for batch inference in your configured Amazon Bedrock regions can be used — the same identifiers as Chat Completions and Embeddings. To shortlist them, call search_models with route=openai_chat_completion&batch=true, or route=openai_embedding&batch=true for embeddings; each entry also carries a batch field.

The shortlist is a hint, not a rule

batch is advertised on a best-effort basis and never used to reject a request. A model it does not advertise — or says nothing about — may still run a batch, so submit the batch rather than ruling the model out; the answer you get back is the authoritative one.

A model that cannot serve batched requests is refused when the batch is created, naming the model. A model this deployment normally serves through another Amazon Bedrock endpoint is batched under the identifier the batch endpoint knows it by, so it needs nothing from you.

A batch's own model field, and the model field of each request in the JSONL body, may name a wildcard pattern instead of an exact model. It is resolved once, when the batch is created: the concrete model that pattern meant that day is what the batch reports and runs for its whole life, even after a newer release would resolve it differently.

Listing order

Batches are listed most recent first, ordered by the created_at each one reports, so paging with the after cursor and sorting a page on created_at give the same sequence. Batches sharing a second are ordered by identifier, so the sequence is stable from one page to the next. created_at is the moment the batch was created, and the 24-hour processing window reported as expires_at runs from it.

The Listing Window

A listing does not reach every batch ever created. It is answered from a window of up to 1,000 of the most recent batch records held in the bucket set by AWS_S3_BUCKET, and an after cursor naming a batch outside that window returns an empty page. Finding that window costs a bounded number of storage requests, so however many batches the bucket holds, the window is taken from the newest of them as long as that seek converges within its budget. Three things narrow it further:

  • A burst larger than one listing's whole budget can outrun the seek, widening the window toward the older end of that burst instead of reaching the newest batch in it. The seek crosses several thousand batch records per listing whatever their density, so this needs a burst larger than that with nothing created since — each batch requires its own inference job, and a real deployment is bound by Bedrock's batch job quota long before it reaches this.
  • The window is shared with the Anthropic Message Batches surface: records created through either API count against the same window, and each listing then shows only its own.
  • Deleted batches keep their slot. A deleted batch is not listed, but its record still occupies one place in the window — and deletion is served by the Anthropic surface only, see Deleting a batch.

Retrieving a batch by its identifier is unaffected — that works for as long as the record exists. Keep the identifiers you need rather than relying on the listing to find them again.

Batches Created Before 1.17

A batch created by an earlier version reports a created_at captured a moment after the batch itself was created, which can place it a few seconds out of order relative to a batch created alongside it. Those values are kept as they were recorded. While such batches can still appear in a listing, page with the after cursor rather than rebuilding the order from created_at.

Prerequisites

The Batch API is disabled until the deployment declares an AWS IAM service role that Amazon Bedrock assumes to read the requests and write the results:

The permissions the role and the server need are listed in IAM Permissions. While the role is unset, every batch endpoint answers 503.

Billing

Batched requests are billed at the published batch rate for the model, roughly half the on-demand rate. Usage is recorded once, when the batch ends. See Cost Management.

Limits and behaviour to know

Caps enforced at creation

Limit Value
Minimum requests per batch 100 (default quota)
Maximum requests per batch 50,000
Maximum input file size 200 MB
custom_id length 64 characters
Distinct models per input file 1 (upstream rule)
Processing window 24 hours from creation

A batch below the minimum, or past any of these caps, is refused when it is created and the message names the shortfall, rather than accepted and failed later.

The 100-request minimum is a quota default

100 is the default of the Amazon Bedrock quota Minimum number of records per batch inference job, which is set per model and adjustable for some of them — see Amazon Bedrock quotas. The gateway checks against that default, not against your account's own value, so a raised quota is enforced by Amazon Bedrock rather than here — a batch of 150 clears this check and is then refused by the backend — and a lowered one is not usable: fewer than 100 requests is still refused here.

What the input file is checked for

The whole file is read and checked while POST /v1/batches is answered. A file that cannot be batched as it stands is refused there: the call returns 400 naming the line at fault, and no batch is created — no request runs, and there is nothing to poll or to cancel.

What the line carries Answer
Not valid JSON, truncated, or a JSON value that is not an object Line 8: each line must be a JSON object.
A url other than the batch's own endpoint Line 4: 'url' must be '/v1/chat/completions', the endpoint the batch targets.
A method other than POST Line 5: 'method' must be 'POST', not 'GET'.
No body, or a body that is not an object Line 9: 'body' must be a JSON object.
An empty custom_id, or one past its length cap Line 7: 'custom_id' must be between 1 and 64 characters.
A custom_id another line already used Line 6: 'custom_id' 'req-4' is used more than once.
A request the endpoint itself would refuse That endpoint's own validation error, naming the field

The file as a whole is read the same way: one uploaded with a purpose other than batch, one holding no request at all, one past the caps above, and one naming more than one model are each refused by the create call.

Coming from the OpenAI API: read the create call, not only the status

Upstream, a file problem is reported after the batch exists: creation returns a batch in validating, which then becomes failed carrying one errors.data[] entry per offending line. Here the same problems are answered by the create call itself, so an SDK raises on client.batches.create(...) — a BadRequestError carrying the message above — and the polling loop that follows it never runs. Catch that exception; the rest of the workflow is unchanged.

failed is still reported, for a batch that was accepted and then could not run at all. Its single errors.data[] entry describes the batch rather than a line, so line is absent, and the batch names no result file.

How many batches may be created at once

POST /v1/batches does its work while the call is open: the input file is read and every request in it is translated, which fetches whatever those requests point at — an image URL, a document, an audio file. That work is bounded so a few large files cannot exhaust the server, and the bound is answered rather than queued.

Bound Value What a caller over it reads
Input files being translated at once, server-wide 2 503, after waiting up to 30 seconds for one of them to finish
Creations in flight for one tenant API key 8 429, immediately

The per-key cap counts tenant API keys only: it is a fairness rule between tenants, and a deployment key is the operator's own traffic, held by the server-wide slots like everything else.

Both are refusals of the create call alone: nothing is read, no batch is created, and no other endpoint is affected. Both are safe to retry, and the message names what to wait for. Neither carries a retry-after header.

A creation that is admitted is never held up by another for longer than one translation wave: the server-wide slots are given back between waves, not at the end of a file, so a 50,000-request batch whose inputs are slow to fetch shares the server with every other caller instead of owning it until it is done. A client that disconnects while its creation waits releases its place.

Creating many batches from one tenant key

Create them a few at a time rather than firing every file at once: a 429 here means the previous creations are still reading their files, so the work is already under way. Upload the files first — POST /v1/files is not bounded this way — then create the batches as each creation answers.

Deleting a batch

This surface serves no DELETE /v1/batches/{batch_id}: a batch record created here keeps its place in the bucket for as long as the bucket holds it. Its result files are ordinary Files API objects, so DELETE /v1/files/{file_id} removes them, and output_expires_after set at creation expires them on their own.

The Anthropic Message Batches API does serve DELETE /v1/messages/batches/{id}, but a batch record is readable only through the API it was created with, so that endpoint deletes Anthropic-created batches only.

Try it

1. Upload the requests

One JSON object per line, uploaded with purpose="batch". Every line names the same model, targets the same endpoint, carries a unique custom_id, and holds the request itself in body.

{"custom_id": "req-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "amazon.nova-micro-v1:0", "messages": [{"role": "user", "content": "Summarize: ..."}]}}
{"custom_id": "req-2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "amazon.nova-micro-v1:0", "messages": [{"role": "user", "content": "Summarize: ..."}]}}

Embedding requests are written the same way, against /v1/embeddings:

{"custom_id": "doc-1", "method": "POST", "url": "/v1/embeddings", "body": {"model": "amazon.titan-embed-text-v2:0", "input": "First passage of the corpus"}}
{"custom_id": "doc-2", "method": "POST", "url": "/v1/embeddings", "body": {"model": "amazon.titan-embed-text-v2:0", "input": "Second passage of the corpus"}}

Both samples above are abridged: a real file needs at least 100 requests, the minimum a batch carries.

from openai import OpenAI

client = OpenAI(base_url="https://your-host/v1", api_key="...")

requests_file = client.files.create(file=open("requests.jsonl", "rb"), purpose="batch")

2. Create the batch

batch = client.batches.create(
    input_file_id=requests_file.id,
    endpoint="/v1/chat/completions",
    completion_window="24h",
)

Example request (curl):

curl -X POST "$BASE/v1/batches" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "input_file_id": "file-06fvfg3lbdqarbad8kbo55g0sg5h3s4a",
    "endpoint": "/v1/chat/completions",
    "completion_window": "24h"
  }'

Example response:

{
  "id": "batch_06fvfg3lbdqarbad8kbo55g0sg5h3s4a",
  "object": "batch",
  "endpoint": "/v1/chat/completions",
  "input_file_id": "file-06fvfg3lbdqarbad8kbo55g0sg5h3s4a",
  "completion_window": "24h",
  "status": "validating",
  "created_at": 1786568013,
  "expires_at": 1786654413,
  "request_counts": {"total": 100, "completed": 0, "failed": 0},
  "model": "amazon.nova-micro-v1:0"
}

3. Poll until it ends

batch = client.batches.retrieve(batch.id)
print(batch.status, batch.request_counts)

4. Read the results

results = client.files.content(batch.output_file_id).text

Each line pairs a custom_id with the completion it produced:

{"id": "batch_req_9f2c...", "custom_id": "req-1", "response": {"status_code": 200, "request_id": "batch_req_9f2c...", "body": {"id": "chatcmpl-req-1", "object": "chat.completion", "choices": [{"index": 0, "message": {"role": "assistant", "content": "..."}, "finish_reason": "stop"}], "usage": {"prompt_tokens": 22, "completion_tokens": 9, "total_tokens": 31}}}, "error": null}

Requests that failed are collected in a separate file, named by error_file_id.

Expiring the result files

Result files are kept until deleted, and are billed as stored objects meanwhile. Pass output_expires_after when creating the batch to have both files expire on their own — between 1 hour and 30 days, counted from the moment they are written:

batch = client.batches.create(
    input_file_id=requests_file.id,
    endpoint="/v1/chat/completions",
    completion_window="24h",
    output_expires_after={"anchor": "created_at", "seconds": 7 * 24 * 3600},
)

Results Are Not in Request Order

Output lines may come back in any order, as upstream also warns. Match a result to its request with custom_id, never with the line number.

Next: Files API · Chat Completions API · Embeddings API · Message Batches API, the only surface with a delete endpoint