Moderations API¶
Classify content for harm with Amazon Bedrock Guardrails, inline guardrail checks (no guardrail resource needed), or Amazon Comprehend toxicity detection through an OpenAI-compatible Moderations interface.
At a glance¶
- Three backends behind one endpoint: your own Bedrock guardrail (categories, thresholds, denied topics, word filters and sensitive-information policies configured in AWS), inline guardrail content filter checks, or Amazon Comprehend toxicity detection.
- Works out of the box. Guardrail checks and Comprehend need no setup at all, so
/v1/moderationsanswers on any deployment, with or without a guardrail configured. - OpenAI model names are aliases.
omni-moderation-latestandtext-moderation-latestresolve to an AWS backend, so an existing integration changes its base URL and nothing else. - Text and images. Every element of an
inputarray is classified independently and gets its own entry inresults; PNG and JPEG images are classified by guardrail-resource models. - Classifications run in your own AWS account — no traffic to third-party endpoints.
- Differs from OpenAI: only the five categories with an AWS counterpart are populated, and scores come back at the backend's own granularity rather than as continuous confidences on every backend.
curl -X POST "$BASE/v1/moderations" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": ["Some text to classify", "Another string to classify"]
}'
Endpoints¶
| Endpoint | Method | What It Does | Powered By | MCP Tool |
|---|---|---|---|---|
/v1/moderations | POST | Classify inputs with a guardrail, guardrail checks, or toxicity detection | Amazon Bedrock Guardrails / Amazon Comprehend | openai_moderation |
Feature compatibility¶
| Feature | Status | Notes |
|---|---|---|
| Input | ||
input string / array of strings | Each element yields one independent result | |
input text parts | {"type": "text", "text": ...} | |
input image parts | Guardrail resource models only (not guardrail checks); PNG and JPEG | |
| Empty string input | Returns an unflagged result without calling AWS (OpenAI parity); whitespace-only strings are not covered by this shortcut and are classified normally (billed Comprehend call) | |
model | Guardrail, Comprehend, or an OpenAI moderation model alias (see below) | |
| Output | ||
flagged | Also raised by guardrail policies without a mapped category, and on Comprehend by the overall toxicity score or unmapped labels such as profanity | |
categories / category_scores | Mapped categories only; OpenAI categories without a counterpart stay false / 0.0 | |
category_applied_input_types | Reflects each classified element's modality | |
| Usage tracking | ||
| Guardrail text units / images | Billing units (one text unit per 1,000 characters per input; one unit per image) | |
| Comprehend units | Billing unit (100 characters, minimum 3 per call) | |
| Other | ||
| Model discovery | Moderation models and their aliases appear in the model listings | |
| Long text inputs | Comprehend inputs of any length are split into API-sized segments transparently |
Legend:
- Supported — Fully compatible with OpenAI API
- Model-Dependent — Behavior depends on the model or backend; check the Notes column
- Partial — Supported with limitations
- Unsupported — Not available in this implementation
- Extra Feature — Enhanced capability beyond OpenAI API
Models¶
The moderation models appear in the /v1/models and /search_models listings (route=openai_moderation) with their OpenAI aliases.
Amazon Bedrock Guardrails¶
| Model | OpenAI aliases | Notes |
|---|---|---|
amazon.bedrock-runtime-guardrail | omni-moderation-latest, omni-moderation-2024-09-26 | The server's default guardrail. Text and image inputs. Listed only when a guardrail is configured |
amazon.bedrock-runtime-guardrail-checks | omni-moderation-* when no guardrail is configured | Inline guardrail content filter checks (InvokeGuardrailChecks) — no guardrail resource needed. Text inputs only. Listed only when a configured Bedrock region offers the operation |
<guardrail-id>, <guardrail-id>:<version>, or guardrail ARN | — | Any explicit guardrail (requires guardrail override to be allowed) |
Amazon Comprehend¶
| Model | OpenAI aliases | Notes |
|---|---|---|
amazon.comprehend-toxicity | text-moderation-latest, text-moderation-stable | Toxicity detection. English-optimized; non-English input is classified on a best-effort basis, no images |
Comparison¶
| Capability | Amazon Bedrock Guardrails | Amazon Bedrock Guardrail Checks | Amazon Comprehend |
|---|---|---|---|
| Setup | Create and configure a guardrail in Amazon Bedrock | None — works out of the box in supported regions | None — works out of the box |
| Text inputs | Any language supported by the guardrail | Languages supported by guardrail content filters | English-optimized; non-English input classified on a best-effort basis |
| Image inputs | PNG and JPEG | Not supported | Not supported |
| Mapped categories | hate, harassment, sexual, violence, illicit | hate, harassment, sexual, violence, illicit | hate, harassment, sexual, violence, violence/graphic |
| Category scores | Quantized confidence levels (0.0 / 0.25 / 0.5 / 0.75) | Severity scores in 0.2 increments (0.0 – 1.0) | Continuous scores (0.0 – 1.0) |
| Custom policies | Denied topics, word filters, PII, prompt attacks, contextual grounding | Content filters only | Fixed toxicity labels |
| Tunable thresholds | Per-filter strengths configured on the guardrail | Fixed flagging threshold (score ≥ 0.5) | Fixed flagging threshold (score ≥ 0.5) |
moderation request parameter | Applied to generations natively | Moderations API only | Moderations API only |
| Input length | ApplyGuardrail text unit limits | InvokeGuardrailChecks text unit limits | Unlimited (split into 1 KB segments transparently) |
Selecting the Model¶
The model parameter selects the moderation model:
model value | Model used |
|---|---|
| Omitted | The server's default guardrail; when none is configured, guardrail checks in a supported region, then Comprehend toxicity detection as a last resort |
amazon.bedrock-runtime-guardrail | The server's default guardrail (an error when none is configured) |
amazon.bedrock-runtime-guardrail-checks | Inline guardrail content filter checks, even when a guardrail is configured (an error when no configured region offers the operation) |
omni-moderation-* | Same as an omitted model |
amazon.comprehend-toxicity / text-moderation-* | Comprehend toxicity detection, even when a guardrail is configured |
<guardrail-id> or <guardrail-id>:<version> | That guardrail (requires guardrail override to be allowed) |
| Guardrail ARN | That guardrail, applied in the region embedded in the ARN |
The server guardrail comes from AWS_BEDROCK_GUARDRAIL_IDENTIFIER / AWS_BEDROCK_GUARDRAIL_VERSION, or from the X-Amzn-Bedrock-GuardrailIdentifier / X-Amzn-Bedrock-GuardrailVersion request headers when AWS_BEDROCK_ALLOW_GUARDRAIL_OVERRIDE is enabled. Explicit guardrails in model also require that setting.
Guardrails are regional: a plain guardrail ID is applied in the primary Bedrock region, while an ARN selects its own region. Comprehend calls use AWS_COMPREHEND_REGION (with multi-region failover otherwise).
Guardrail checks (InvokeGuardrailChecks) are available in a limited set of AWS regions only (currently us-east-1, us-east-2, us-west-2, eu-west-2, eu-north-1, ap-northeast-1, and ap-southeast-2): calls run in the configured Bedrock regions that offer the operation, in priority order with multi-region failover, and the backend is unavailable when none of them does.
Moderating generations directly
The guardrail selection and category mapping also power the moderation request parameter of the Chat Completions and Responses APIs: the guardrail is applied to the generation itself, and the classification of the input and output is reported in the response's moderation field — for Chat Completions on non-streaming requests only, and for Responses also on the terminal event when streaming. The moderation parameter requires a guardrail — Comprehend is not available there — and is rejected (400) on Amazon Bedrock Mantle-served models.
Category Mapping¶
Amazon Bedrock Guardrails — content policy filters map to the OpenAI moderation categories:
| Amazon Bedrock filter | OpenAI category |
|---|---|
HATE | hate |
INSULTS | harassment |
SEXUAL | sexual |
VIOLENCE | violence |
MISCONDUCT | illicit |
Filter confidence levels become scores: NONE → 0.0, LOW → 0.25, MEDIUM → 0.5, HIGH → 0.75.
Every other guardrail policy — denied topics, word filters, sensitive information (PII), prompt attacks, contextual grounding — still contributes to the top-level flagged field whenever the guardrail intervenes, even though no individual category is set.
Amazon Bedrock Guardrail Checks — the inline contentFilter check uses the same five-category mapping as guardrail content policy filters, but reports each category's severity score (0.0 to 1.0 in 0.2 increments) directly in category_scores. A category — and the input — is flagged when its score reaches 0.5.
Amazon Comprehend — toxicity labels map to the OpenAI moderation categories, with their detection scores (0.0–1.0) reported directly:
| Comprehend label | OpenAI category |
|---|---|
HATE_SPEECH | hate |
HARASSMENT_OR_ABUSE | harassment |
INSULT | harassment |
SEXUAL | sexual |
VIOLENCE_OR_THREAT | violence |
GRAPHIC | violence/graphic |
PROFANITY | (flagged only) |
An input is flagged when its overall toxicity or any label score reaches 0.5. Long texts are split into 1 KB segments and the highest score per category is kept. Comprehend's DetectToxicContent API is always called with the English language code, regardless of the input's actual language.
With either model, OpenAI sub-categories without a counterpart (e.g. self-harm, sexual/minors) are always false.
Working with the Moderations API¶
Inputs¶
Each input element is classified independently and yields one entry in results:
inputas a string — one text classification.inputas an array of strings — one classification per string.inputas an array of parts —{"type": "text", "text": ...}and{"type": "image_url", "image_url": {"url": ...}}parts. Images must be PNG or JPEG, and require a guardrail resource model (guardrail checks and Comprehend are text-only).
Each result's category_applied_input_types reflects the classified element's modality: ["text"] for every category on text inputs; on image inputs, ["image"] for the categories that support images and [] for the text-only ones.
MCP / AI agent usage: image_url.url accepts an HTTPS URL, data URI (data:<mime>;base64,<data>), base64 string, or S3 URI — no binary upload needed.
Billing¶
-
Guardrails — AWS bills per text unit and per image processed by the ApplyGuardrail API, once for every policy the guardrail applies, each at that policy's own rate; see Amazon Bedrock pricing. No Bedrock model invocation is involved. The units AWS reports per policy appear in usage logs and cost tracking as
text_unitsandinput_images, under one model per applied policy:Model Policy amazon.bedrock-runtime-guardrail-contentContent filters ( text_units, andinput_imagesfor image content)amazon.bedrock-runtime-guardrail-topicDenied topics amazon.bedrock-runtime-guardrail-wordWord filters amazon.bedrock-runtime-guardrail-sensitive-informationSensitive information (paid entity types) amazon.bedrock-runtime-guardrail-sensitive-information-freeSensitive information (free entity types) amazon.bedrock-runtime-guardrail-contextual-groundingContextual grounding amazon.bedrock-runtime-guardrail-automated-reasoningAutomated reasoning A guardrail applying several policies therefore produces several usage entries for one request, and its cost is their sum. - Guardrail checks — AWS bills per text unit and per requested check; the gateway requests the
contentFiltercheck only. Billed text units appear in usage logs and cost tracking astext_unitsunder theamazon.bedrock-runtime-guardrail-checksmodel. - Comprehend — AWS bills toxicity detection per 100-character unit with a 3-unit minimum per call; see Amazon Comprehend pricing. Billed units appear in usage logs and cost tracking ascomprehend_unitsunder theamazon.comprehend-toxicitymodel.
Limits and behaviour to know¶
An input array holds at most 2048 elements. Each element is classified — and billed — independently by a separate AWS call, so large arrays incur a proportional number of AWS calls.
model names an exact value from Selecting the Model: this route chooses the classification model before the request is examined, so a wildcard pattern is rejected, unlike on the routes that generate a response.
Guardrail checks moderation calls the bedrock:InvokeGuardrailChecks IAM action; see IAM Permissions. A deployment without that permission keeps working: where guardrail checks are only reached as the default omni-moderation-* resolution, an AccessDenied response degrades the request to Comprehend toxicity detection and records a warning in the log. A request that names amazon.bedrock-runtime-guardrail-checks explicitly is not degraded.
Request headers¶
The Amazon Bedrock guardrail headers select the guardrail this endpoint classifies with. All headers are optional.
Content Safety (Guardrails)¶
| Header | Purpose | Valid Values |
|---|---|---|
X-Amzn-Bedrock-GuardrailIdentifier | Guardrail ID for content filtering | Your guardrail identifier |
X-Amzn-Bedrock-GuardrailVersion | Guardrail version | Version number (e.g., 1) |
Both headers are honoured only when AWS_BEDROCK_ALLOW_GUARDRAIL_OVERRIDE is enabled — otherwise the server guardrail applies, as described under Selecting the model. X-Amzn-Bedrock-Trace is accepted but has no effect on this route — categories and scores are read from the guardrail's assessments, which are always returned.
Example with headers:
curl -X POST "$BASE/v1/moderations" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-H "X-Amzn-Bedrock-GuardrailIdentifier: your-guardrail-id" \
-H "X-Amzn-Bedrock-GuardrailVersion: 1" \
-d '{
"input": "Some text to classify"
}'
No performance headers on this route
X-Amzn-Bedrock-Service-Tier and X-Amzn-Bedrock-PerformanceConfig-Latency have no effect here: classification calls ApplyGuardrail, guardrail checks or Amazon Comprehend, none of which invokes a Bedrock model.
Detailed Documentation
For complete information about these headers, configuration options, and use cases, see:
Try it¶
Classify one string:
curl -X POST "$BASE/v1/moderations" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": "Some text to classify"
}'
Response:
{
"id": "modr-0f1b3c6e8d9a4b5c",
"model": "abcd1234efgh:1",
"results": [
{
"flagged": true,
"categories": {"hate": true, "harassment": false, "...": false},
"category_scores": {"hate": 0.75, "harassment": 0.25, "...": 0.0},
"category_applied_input_types": {"hate": ["text"], "harassment": ["text"], "...": ["text"]}
}
]
}
Classify text and an image together:
curl -X POST "$BASE/v1/moderations" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": [
{"type": "text", "text": "Describe this image"},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.png"}}
]
}'
Next steps¶
Next: Guardrails configuration · Chat Completions API · Responses API · Model pricing