Skip to content

Configuration Guide

stdapi.ai is configured entirely through environment variables, which are read once at startup and cannot be changed without restarting the service. This guide explains each setting category with practical examples to help you configure the service correctly.

What you can configure:

  • AWS regions - Access models across multiple regions for availability and model selection
  • Data sovereignty - Control which AWS regions are used for compliance (GDPR, HIPAA, etc.)
  • Storage - S3 buckets for file operations, regional buckets for multi-region deployments
  • Authentication - API keys via SSM or Secrets Manager for secure access control
  • Observability - Logging levels, OpenTelemetry, request/response debugging
  • Security - CORS, proxy headers, trusted hosts for production deployments
  • Performance - Caching, model overrides, S3 acceleration
  • TLS / SSL - End-to-end encryption using Granian environment variables

Zero Configuration Startup

stdapi.ai works out of the box with zero configuration. The service automatically detects your current AWS region and discovers available Bedrock models.

Prerequisites

Before configuring stdapi.ai, ensure you have:

  • AWS Account with access to Amazon Bedrock
  • AWS Credentials configured via environment variables, AWS CLI, or IAM role (for EC2/ECS/Lambda deployments)
  • IAM Permissions to access required AWS services (see the IAM Permissions guide)
  • S3 Bucket (optional, but recommended for production use with file operations)

Container Runtime

Both the AWS Marketplace and community Docker images run using Granian, a high-performance Python ASGI server. In addition to the stdapi.ai-specific configuration variables documented below, you can also use Granian environment variables to configure the server runtime (e.g., GRANIAN_PORT, GRANIAN_WORKERS, GRANIAN_THREADS, etc.).

The images listen on IPv4 only (GRANIAN_HOST=0.0.0.0). Set GRANIAN_HOST=:: to bind a dual-stack socket answering both IPv4 and IPv6 clients. This is needed wherever a client may resolve the server to an IPv6 address — in particular with ECS service discovery, which publishes an AAAA record for every task in an IPv6-enabled subnet, and some clients (Node.js among them) try that address first and fail with ECONNREFUSED against an IPv4-only listener. The official Terraform module sets it for you when the VPC has IPv6 enabled.

Quick Start

For production deployments, configure these essential settings:

Minimal Production Setup

Single-region deployment with file storage only.

# S3 bucket for file storage (must be in same region as your server)
export AWS_S3_BUCKET=my-stdapi-bucket

# AWS_BEDROCK_REGIONS is optional - will auto-detect your current AWS region if not specified

Production with Authentication

Adds secure API key authentication via AWS Systems Manager.

# S3 bucket for file storage (must be in same region as your server)
export AWS_S3_BUCKET=my-stdapi-bucket

# Secure API authentication (recommended: SSM Parameter Store)
export API_KEY_SSM_PARAMETER=/stdapi/prod/api-key

# AWS_BEDROCK_REGIONS is optional - will auto-detect your current AWS region if not specified

Full Production Setup (All Features Enabled)

Multi-region deployment with all AWS AI services, observability, and security features.

# Core AWS configuration - host server in first region
export AWS_BEDROCK_REGIONS=us-east-1,us-west-2,eu-west-1

# S3 bucket for file storage (must be in us-east-1, your first/primary region)
export AWS_S3_BUCKET=my-stdapi-us-east-1-bucket

# Optional: Transcribe S3 bucket (defaults to AWS_S3_BUCKET if not specified)
# Only set this if you need a separate bucket or if transcribe is in a different region
# export AWS_TRANSCRIBE_S3_BUCKET=my-stdapi-transcribe-us-east-1

# Optional: Regional buckets for async/batch inference in other regions
export AWS_S3_REGIONAL_BUCKETS='{"us-west-2": "my-stdapi-us-west-2-bucket", "eu-west-1": "my-stdapi-eu-west-1-bucket"}'

# AWS AI services regions (optional - when unset, every AWS_BEDROCK_REGIONS entry is a
# candidate with automatic failover; set one to pin the service to a single region)
export AWS_POLLY_REGION=us-east-1           # Text-to-speech
export AWS_TRANSCRIBE_REGION=us-east-1      # Speech-to-text (audio transcription)
export AWS_COMPREHEND_REGION=us-east-1      # Language detection & moderation
export AWS_TRANSLATE_REGION=us-east-1       # Text translation

# Authentication
export API_KEY_SSM_PARAMETER=/stdapi/prod/api-key

# Logging
export LOG_LEVEL=warning
export LOG_CLIENT_IP=true

# Optional: OpenTelemetry observability (AWS X-Ray integration)
# export OTEL_ENABLED=true
# export OTEL_SERVICE_NAME=stdapi-production
# export OTEL_SAMPLE_RATE=0.1

# Production security settings (when behind AWS ALB/CloudFront)
export ENABLE_PROXY_HEADERS=true

# Note: TRUSTED_HOSTS not recommended with AWS ALB - use ALB host-based routing instead
# Only use TRUSTED_HOSTS if you cannot configure host validation at the load balancer level

# Optional: CORS for browser-based web applications
# export CORS_ALLOW_ORIGINS='["https://app.example.com"]'

Development Setup

Local development configuration with API documentation and debug logging enabled.

# Minimal configuration for local development
export AWS_S3_BUCKET=my-stdapi-dev-bucket

# Enable API documentation
export ENABLE_DOCS=true
export ENABLE_REDOC=true

# Full request/response logging for debugging
export LOG_LEVEL=info
export LOG_REQUEST_PARAMS=true

# AWS_BEDROCK_REGIONS is optional - will auto-detect your current AWS region if not specified

S3 Bucket Required for Certain Features

Without an S3 bucket configured, some features will be disabled (such as image output as URL, audio transcription). See the relevant API documentation for feature requirements.

All Other Settings Are Optional

The configurations above are sufficient for most production deployments. All other settings can be configured as needed for your specific use case.

Environment Variable Summary

This section provides a quick reference of all available configuration options. Detailed explanations for each variable can be found in the sections below.

Essential (Production)

Variable Default Description
AWS_S3_BUCKET None Primary S3 bucket for file storage; must be in first region of AWS_BEDROCK_REGIONS
AWS_BEDROCK_REGIONS Current region Comma-separated regions for Bedrock; first region is where server should be hosted

AWS Client

Variable Default Description
AWS_ADAPTIVE_RETRY false Enable adaptive retry mode that throttles back under congestion rather than using fixed exponential backoff
AWS_MAX_POOL_CONNECTIONS 50 Maximum concurrent HTTP connections per AWS service client
AWS_CONNECT_TIMEOUT 5 Timeout in seconds for establishing a connection to an AWS service endpoint, and for a real-time audio session to become ready in a region

AWS Storage

Variable Default Description
AWS_S3_ACCELERATE false Enable S3 Transfer Acceleration for faster global downloads via CloudFront edge locations
AWS_S3_REGIONAL_BUCKETS {} Region-specific S3 buckets for Bedrock async/batch inference operations
AWS_S3_ACCEPTED_BUCKETS {} External S3 buckets with read access, mapped to their region for S3 URI conversion and routing
AWS_S3_TMP_PREFIX tmp/ S3 prefix for temporary files used for jobs; configure lifecycle policies on this prefix
AWS_S3_FILES_PREFIX files/ S3 prefix for Files API objects; configure S3 lifecycle policies on this prefix
AWS_S3_VIDEOS_PREFIX videos/ S3 prefix for generated videos (Videos API); persists until deleted through the API
AWS_S3_VIDEOS_EXPIRES_AFTER None Retention period in seconds for generated videos; sets Video.expires_at and blocks expired downloads
AWS_S3_BATCHES_PREFIX batches/ S3 prefix for Batch API data (requests, results, batch records); configure lifecycle policies on it
AWS_S3_VECTORS_BUCKET None Amazon S3 vector bucket backing the Vector Stores API; unset disables it
AWS_S3_VECTORS_REGION First Bedrock region Region holding AWS_S3_VECTORS_BUCKET; the vector bucket has no failover
AWS_S3_VECTOR_STORES_PREFIX vector_stores/ S3 prefix for the Vector Stores API records (stores, attached files, batches)
VECTOR_STORE_EMBEDDING_MODEL amazon.titan-embed-text-v2:0 Model embedding the indexed files and the search queries
VECTOR_STORE_CHUNK_SIZE_TOKENS 800 Default chunk size for files indexed without an explicit chunking_strategy
VECTOR_STORE_CHUNK_OVERLAP_TOKENS 400 Default chunk overlap; must not exceed half the chunk size
AWS_SQS_VECTOR_STORE_QUEUE_URL None Amazon SQS queue making vector store indexing survive the server running it; unset keeps it in-process
AWS_BEDROCK_KNOWLEDGE_BASE_IDS [] Allowlist of Amazon Bedrock knowledge bases addressed as vs_kb_... vector stores; empty disables it
AWS_TRANSCRIBE_S3_BUCKET AWS_S3_BUCKET S3 bucket for temporary audio transcription files; must be in same region as AWS_TRANSCRIBE_REGION
AWS_TRANSCRIBE_OUTPUT_ENCRYPTION_KEY_ARN None AWS KMS key encrypting the transcription output objects; unset keeps the bucket's own encryption
AWS_TRANSCRIBE_STREAM_LANGUAGES [] Languages a streamed transcription picks between when the request names none

AWS AI Services

Variable Default Description
AWS_POLLY_REGION All AWS_BEDROCK_REGIONS Region for Amazon Polly; unset = per-engine regional discovery with automatic failover
AWS_COMPREHEND_REGION All AWS_BEDROCK_REGIONS Region for Amazon Comprehend (language detection, toxicity moderation); unset = automatic failover across all Bedrock regions
AWS_TRANSCRIBE_REGION All AWS_BEDROCK_REGIONS Region for Amazon Transcribe; unset = failover across Bedrock regions with a co-located bucket
AWS_TRANSLATE_REGION All AWS_BEDROCK_REGIONS Region for Amazon Translate; unset = automatic failover across all Bedrock regions

Resilience & Failover

Variable Default Description
AWS_BEDROCK_REGION_ROUTING ordered Region routing strategy: disabled, ordered, lowest_latency, or round_robin (details)
AWS_BEDROCK_REGION_ROUTING_QUOTA_BACKOFF_SECONDS 60 Base interval in seconds for exponential quota backoff per region
AWS_BEDROCK_REGION_ROUTING_MAX_QUOTA_BACKOFF_SECONDS 3600 Hard ceiling in seconds on the exponential quota backoff per region (default: 1 hour)
AWS_BEDROCK_REGION_ROUTING_QUOTA_STALE_FACTOR 2 Multiplier on max quota backoff to determine when the consecutive-error counter resets
AWS_BEDROCK_REGION_ROUTING_UNAVAILABLE_BACKOFF_SECONDS 30 Seconds to avoid a region after unavailability errors
AWS_BEDROCK_MAX_RETRIES 9 Cap on the retries per Bedrock invocation; with region routing, each candidate region is tried at most once
AWS_FAILOVER_MAX_RETRIES 2 SDK retries per candidate region for the multi-region failover services (Polly, Transcribe, Translate, Comprehend)

Bedrock Mantle

Variable Default Description
AWS_BEDROCK_MANTLE_ENABLED true Expose models served by the Amazon Bedrock Mantle endpoint alongside classic Bedrock Converse models
AWS_BEDROCK_MANTLE_REGIONS Mantle-capable subset of AWS_BEDROCK_REGIONS AWS regions used for Bedrock Mantle, in failover priority order
AWS_BEDROCK_MANTLE_ENDPOINT_URL None Override the Bedrock Mantle endpoint URL template ({region} placeholder)
AWS_BEDROCK_MANTLE_PREFERRED_MODELS [] Model IDs served via Mantle even when also available on the classic bedrock-runtime endpoint
AWS_BEDROCK_MANTLE_SERVICE_HEADER false Honor the x-stdapi-service: bedrock-mantle request header to route dual-homed models through Mantle per request
AWS_BEDROCK_MANTLE_PROJECT None Default Bedrock Project/Workspace ID applied to Mantle requests for cost tracking and observability
AWS_BEDROCK_ALLOW_MANTLE_PROJECT_OVERRIDE false Allow requests to override the configured Mantle project via the OpenAI-Project / anthropic-workspace header
AWS_BEDROCK_EXTERNAL_WEB_ACCESS false Let the built-in web search tool reach the public web instead of the Amazon Bedrock web index
AWS_BEDROCK_ALLOW_EXTERNAL_WEB_ACCESS_OVERRIDE false Allow requests to override external web access with the external_web_access extra model parameter

Bedrock Advanced

Variable Default Description
AWS_BEDROCK_CROSS_REGION_INFERENCE true Allow automatic model routing to other configured regions
AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL true Allow global cross-region inference routing to any region worldwide (disable for GDPR compliance)
AWS_BEDROCK_MODEL_REGION_RESTRICT {} Restrict a model to specific region(s) only (e.g. for region-specific features like Nova grounding)
AWS_BEDROCK_LEGACY false Allow usage of deprecated/legacy Bedrock models
AWS_BEDROCK_DEPRECATED_MODEL_FALLBACK true Transparently reroute requests using a deprecated model ID to its recommended replacement
AWS_BEDROCK_DEPRECATED_MODELS {} Additional deprecated model mappings merged with the built-in registry at startup
AWS_BEDROCK_MARKETPLACE_AUTO_SUBSCRIBE true Allow automatic subscription to new models in AWS Marketplace
AWS_BEDROCK_ALLOW_CROSS_REGION_INFERENCE_PROFILE_ARN false Allow users to pass cross-region inference profile ARNs directly as model IDs
AWS_BEDROCK_ALLOW_APPLICATION_INFERENCE_PROFILE_ARN false Allow users to pass application inference profile ARNs directly as model IDs
AWS_BEDROCK_ALLOW_PROMPT_ROUTER_ARN false Allow users to pass prompt router ARNs directly as model IDs
AWS_BEDROCK_ALLOW_PROMPT_ARN false Allow users to reference Prompt Management prompt ARNs in the Responses API prompt parameter
AWS_BEDROCK_MODEL_ARN_MAPPING {} Map model IDs to custom inference profile or prompt router ARNs (server-controlled routing)
AWS_BEDROCK_GUARDRAIL_IDENTIFIER None Bedrock Guardrails ID for content filtering and safety controls
AWS_BEDROCK_GUARDRAIL_VERSION None Bedrock Guardrails version number (required with identifier)
AWS_BEDROCK_GUARDRAIL_TRACE None Guardrails trace level: disabled, enabled, or enabled_full
AWS_BEDROCK_ALLOW_GUARDRAIL_OVERRIDE false Allow users to override global guardrail configuration via request headers (security: default off)
AWS_BEDROCK_SESSION_ENCRYPTION_KEY_ARN None KMS key ARN encrypting Amazon Bedrock session storage (Responses API store=true)
AWS_BEDROCK_BATCH_ROLE_ARN None Service role Amazon Bedrock assumes to run batch inference jobs; unset disables the Batch APIs
AWS_BEDROCK_USER_ROLE_ARN None Run each end user's model calls under a role session of their own, so AWS reports their spend separately
AWS_BEDROCK_USER_ROLE_SESSION_DURATION 3600 Lifetime in seconds of a per-end-user role session (900–3600)
AWS_BEDROCK_USER_ROLE_TAG_KEY user Session tag key carrying the end user identity, for cost allocation and access policies
AWS_BEDROCK_USER_ROLE_REQUIRE_IDENTITY false Reject a model request that identifies no end user instead of billing it to the server

Authentication

Configure one API key source. If several are set, precedence is API_KEY → SSM Parameter Store → Secrets Manager — see Authentication:

Variable Default Description
API_KEY_SSM_PARAMETER None AWS Systems Manager Parameter Store path for API key (recommended)
API_KEY_SECRETSMANAGER_SECRET None AWS Secrets Manager secret name containing API key
API_KEY_SECRETSMANAGER_KEY api_key JSON key name within Secrets Manager secret
API_KEY None Direct API key value (not recommended for production)
AUTHENTICATION_MODE any Accepted methods: any, api_key or cognito

Amazon Cognito user pool tokens are an alternative to the API key — see Amazon Cognito Authentication:

Variable Default Description
AWS_COGNITO_USER_POOL_ID None User pool whose tokens authenticate clients (enables the method)
AWS_COGNITO_CLIENT_IDS None App client IDs whose tokens are accepted (required with a pool)
AWS_COGNITO_REQUIRED_SCOPES None Scopes a token must all carry
AWS_COGNITO_ACCEPT_ID_TOKEN false Also accept identity tokens, not only access tokens
AWS_COGNITO_ISSUER_TYPE original Pool issuer configuration: original or updated

Publishing where tokens come from lets an AI agent authenticate itself — see Authentication Discovery:

Variable Default Description
OAUTH_RESOURCE_IDENTIFIER None Public URL clients dial (publishes the discovery document)
OAUTH_AUTHORIZATION_SERVERS The user pool issuer Issuer URLs of the authorization servers
OAUTH_SCOPES_SUPPORTED Required scopes Scopes a token needs, advertised to clients

API Compatibility

Variable Default Description
OPENAI_ROUTES_PREFIX None (root) Base path prefix for OpenAI-compatible API routes
ANTHROPIC_ROUTES_PREFIX /anthropic Base path prefix for Anthropic-compatible API routes
COHERE_ROUTES_PREFIX /cohere Base path prefix for Cohere-compatible API routes

Logging

Variable Default Description
LOG_LEVEL info Minimum log severity: info, warning, error, critical, or disabled
LOG_REQUEST_PARAMS false Include request/response parameters in logs (not recommended for production)
LOG_CLIENT_IP false Log client IP addresses (requires ENABLE_PROXY_HEADERS for real IPs behind proxies)

CloudWatch Metrics

Variable Default Description
CLOUDWATCH_METRICS false Emit per-request AWS-billed usage as CloudWatch EMF log lines
CLOUDWATCH_METRICS_NAMESPACE stdapi CloudWatch namespace for the emitted usage metrics

Cost Tracking

Variable Default Description
COST_TRACKING false Enable real-time cost computation from live AWS pricing
COST_PRICE_OVERRIDES {} JSON map of operator-supplied unit prices for models missing from the AWS catalog

Observability (OpenTelemetry)

Variable Default Description
OTEL_ENABLED false Enable distributed tracing via OpenTelemetry (integrates with AWS X-Ray, Jaeger, etc.)
OTEL_SERVICE_NAME stdapi.ai Service name identifier in trace visualizations
OTEL_EXPORTER_ENDPOINT http://127.0.0.1:4318/v1/traces OTLP HTTP endpoint URL for trace export
OTEL_SAMPLE_RATE 1.0 Trace sampling rate from 0.0 (none) to 1.0 (all requests)

HTTP/Security

Variable Default Description
CORS_ALLOW_ORIGINS None JSON array of allowed origins for browser cross-origin requests
TRUSTED_HOSTS None JSON array of trusted Host header values (prefer ALB host-based routing; see details)
ENABLE_PROXY_HEADERS false Trust X-Forwarded-* headers from reverse proxies (only enable behind trusted proxy)
PROXY_TRUSTED_HOSTS * Peer IPs/ranges whose X-Forwarded-* headers are trusted (restrict from * for safety)
GRANIAN_HOST 0.0.0.0 Listener bind address; :: binds a dual-stack socket answering IPv4 and IPv6 clients
GRANIAN_SSL_CERTIFICATE None Path to SSL certificate file for end-to-end encryption
GRANIAN_SSL_KEYFILE None Path to SSL private key file (PKCS#8) for end-to-end encryption
GRANIAN_SSL_KEYFILE_PASSWORD None Password for the SSL private key file
GRANIAN_SSL_PROTOCOL_MIN tls1.3 Minimum supported TLS version (tls1.2 or tls1.3)
GRANIAN_SSL_CA None Path to CA certificate bundle for client verification (mTLS)
GRANIAN_SSL_CLIENT_VERIFY false Enable client certificate verification (mTLS)
ENABLE_GZIP false Enable GZip compression for responses >1KB (prefer AWS ALB/CloudFront compression)
SSRF_PROTECTION_BLOCK_PRIVATE_NETWORKS true Block requests to private/local networks for SSRF protection
MAX_INPUT_FILE_SIZE 0 Maximum size in bytes of an inline input file loaded into memory (0 disables)
MAX_CONCURRENT_INPUT_DOWNLOADS 8 Maximum input files fetched/resolved concurrently per request

Application Behavior

Variable Default Description
TIMEZONE UTC IANA timezone identifier for request timestamps
STRICT_INPUT_VALIDATION false Reject API requests with unknown/extra fields
CHAT_COMPLETIONS_REASONING_FIELD reasoning_content Field carrying reasoning text on /v1/chat/completions: reasoning_content, reasoning, or none
MODEL_ALIASES {} JSON object mapping custom model name aliases to Bedrock model IDs, optionally with per-alias configuration
DEFAULT_TTS_MODEL amazon.polly-standard Default TTS model: amazon.polly-standard, -neural, -long-form, or -generative
DEFAULT_TTS_LANGUAGE None Default language for TTS (e.g., en-US); when set, skips Amazon Comprehend auto-detection
TOKENS_ESTIMATION false Deprecated and ignored (token estimation removed)
TOKENS_ESTIMATION_DEFAULT_ENCODING None Deprecated and ignored (token estimation removed)
DEFAULT_MODEL_PARAMS {} JSON object with per-model default inference parameters (temperature, max_tokens, etc.)
DEFAULT_MODEL_SERVICE_TIERS {} JSON object with per-model default service tiers (default, flex, priority, reserved)
AWS_BEDROCK_ALLOW_SERVICE_TIER_OVERRIDE true Allow users to select the service tier per request, overriding the configured one (cost control)
MODEL_CACHE_SECONDS 900 Model list cache lifetime in seconds before lazy refresh (default: 15 minutes)
AI_RESPONSE_TIMEOUT 600 Maximum seconds without data from a model before the request times out (default: 10 min)
SHUTDOWN_DRAIN_TIMEOUT 10 Maximum seconds the server waits for background work to finish after being asked to stop
DROP_UNSUPPORTED_SYSTEM_PROMPT true Drop system prompts for unsupported models; when false, return error instead
ANTHROPIC_BETA_FILTER true Enable filtering of unsupported anthropic_beta flags for Claude models
ANTHROPIC_BETA_ALLOWLIST None Additional anthropic_beta flags to allow beyond built-in Bedrock defaults
EXTRA_MODEL_PARAMS_DENYLIST None Additional "extra model parameters" names to strip, beyond the built-in LiteLLM control-parameter denylist
EXTRA_MODEL_PARAMS_DROP_ALL false Disable the "extra model parameters" passthrough entirely
IMAGE_GENERATION_MODEL None Default Bedrock image model ID used when the image_generation Responses API tool is invoked
REALTIME_CLIENT_SECRET_KEY None Secret the Realtime API's ephemeral client secrets are signed with; derived from the API key when unset
REALTIME_ALLOW_SESSION_OVERRIDE true Allow a client holding an ephemeral client secret to override the session configuration it carries

API Documentation

Variable Default Description
ENABLE_DOCS false Enable interactive Swagger UI documentation at /docs
ENABLE_REDOC false Enable ReDoc documentation UI at /redoc
ENABLE_OPENAPI_JSON false Enable OpenAPI schema endpoint at /openapi.json (auto-enabled with docs/redoc)

MCP (Model Context Protocol)

Variable Default Description
ENABLE_MCP_STREAMABLE_HTTP false Enable MCP server via Streamable HTTP at /mcp — recommended transport
MCP_STATELESS_HTTP false Serve /mcp without server-side sessions — any replica may serve any request
ENABLE_MCP_SSE false Enable MCP server via Server-Sent Events at /sse — legacy transport for older clients
MCP_INCLUDE_TOOLS None Comma-separated tool names to expose exclusively; all others are hidden
MCP_EXCLUDE_TOOLS None Comma-separated tool names to hide; all others remain exposed

AWS Services and Regions

General Configuration

AWS_ADAPTIVE_RETRY

Purpose : Enable adaptive retry mode that adjusts retry pacing based on observed error rates across all AWS service calls

Type : Boolean (true / false)

Default : false

Behavior : When enabled, the retry strategy dynamically responds to real-time congestion signals. If errors are occurring frequently, retries are spaced further apart to avoid amplifying load on an already-stressed endpoint. Once conditions improve, the pacing returns to normal. When disabled, retries follow a standard exponential backoff strategy with fixed intervals. Applies to all AWS services (Bedrock, S3, Polly, Transcribe, etc.).

Latency Impact

Adaptive retry paces retries based on real-time error signals, reducing the risk of retry storms when many clients share the same endpoint under sustained congestion — at the cost of increased per-request latency when throttling is detected, since the client intentionally delays retries to shed load. Prefer it under sustained high load; keep the default standard mode for latency-sensitive, low-traffic workloads.

# Default: standard exponential backoff
export AWS_ADAPTIVE_RETRY=false

# Enable adaptive retry (recommended under sustained high load)
export AWS_ADAPTIVE_RETRY=true

AWS_MAX_POOL_CONNECTIONS

Purpose : Maximum number of concurrent HTTP connections per AWS service client

Type : Integer (must be > 0)

Default : 50

Behavior : Each AWS service client (one per service per region) maintains its own connection pool up to this limit. Under high concurrency, increasing this value prevents requests from queuing for an available connection. Setting it too high may exhaust system file descriptors.

# Default
export AWS_MAX_POOL_CONNECTIONS=50

# High-concurrency deployment
export AWS_MAX_POOL_CONNECTIONS=100

AWS_CONNECT_TIMEOUT

Purpose : Timeout in seconds for establishing a connection to an AWS service endpoint

Type : Integer (must be > 0)

Default : 5

Behavior : Limits how long the client waits when opening a new connection. It also bounds, per candidate region, the time a real-time audio session (generative-voice speech, speech-to-speech) may take to become ready: connection, initial handshake and the first response together. A short value allows fast failover to another region when an endpoint is unreachable. Increase it only if you see spurious connection timeouts on high-latency networks, or real-time audio requests failing with a 503 a few seconds after they start.

# Default: 5 seconds
export AWS_CONNECT_TIMEOUT=5

# High-latency network
export AWS_CONNECT_TIMEOUT=10

Storage Configuration

AWS_S3_BUCKET

Purpose : Primary S3 bucket for storing generated files (images, audio, documents) and temporary data during processing

Default : None (must be configured for file operations)

Best Practice : The bucket must be in the first region specified in AWS_BEDROCK_REGIONS (your primary region where the server should be hosted) to avoid cross-region data transfer costs and reduce latency

export AWS_S3_BUCKET=my-llm-storage-us-east-1

Presigned URLs

Files are served via presigned URLs for secure, time-limited access. Presigned URLs expire after 1 hour.

Terraform Module

When using the Terraform module, the main S3 bucket is created automatically — no manual configuration required.

Startup Warning

If not set, a warning is logged at startup and features that require file storage (image generation, audio output, document processing) will be unavailable.

AWS_S3_ACCELERATE

Purpose : Enable S3 Transfer Acceleration for presigned URLs to improve download performance for large files

Type : Boolean

Default : false

Best Practice : Enable when serving large files (high-resolution images, audio) to geographically distributed users

export AWS_S3_ACCELERATE=true

What is S3 Transfer Acceleration?

S3 Transfer Acceleration uses Amazon CloudFront's globally distributed edge locations to accelerate uploads and downloads to S3 buckets. When enabled, data is routed to the nearest edge location and then transferred to S3 over Amazon's optimized network paths.

Performance Benefits:

  • Faster downloads for users far from your bucket's region
  • Global reach via CloudFront edge locations
  • Optimized routing over Amazon's private backbone network
  • Consistent performance regardless of user location

Typical speed improvements: 50-500% faster for users located far from the bucket region.

Requirements

  1. Enable Transfer Acceleration on your S3 bucket before setting this option:
    aws s3api put-bucket-accelerate-configuration \
      --bucket my-stdapi-bucket \
      --accelerate-configuration Status=Enabled
    
  2. Additional costs: Transfer Acceleration incurs extra data transfer fees. See Amazon S3 Transfer Acceleration pricing

When to Enable

Consider enabling S3 Transfer Acceleration when:

  • Serving generated images via Images API
  • Users are geographically distributed across multiple continents
  • Generating high-resolution images that are large in file size
  • Download performance is critical to user experience

For small images or users close to your bucket region, the performance benefit may not justify the additional cost.

Current Usage

Presigned URLs with Transfer Acceleration are currently only used for the Images API when returning generated images as URLs.

AWS_S3_TMP_PREFIX

Purpose : S3 prefix (folder path) for temporary files used during job processing

Default : tmp/

Best Practice : Configure S3 lifecycle policies to automatically delete objects under this prefix after 1 day

export AWS_S3_TMP_PREFIX=tmp/

What is an S3 Prefix?

An S3 prefix is essentially a folder path within your S3 bucket. When you set AWS_S3_TMP_PREFIX=tmp/, all temporary files are stored under the tmp/ folder structure in your bucket.

Example file paths:

  • With prefix tmp/: s3://my-bucket/tmp/request-id-123/output.json
  • With prefix temporary/: s3://my-bucket/temporary/request-id-123/output.json
  • With empty prefix `:s3://my-bucket/request-id-123/output.json` (not recommended)

Why Use a Prefix?

Using a dedicated prefix for temporary files provides several benefits:

  • Easy Lifecycle Management - Apply S3 lifecycle policies to automatically delete only temporary files
  • Better Organization - Keep temporary files separate from permanent storage
  • Security - Apply different IAM policies or bucket policies to the prefix
  • Cost Control - Easily identify and monitor temporary storage costs

Trailing Slash

Always include a trailing slash (/) in your prefix to create a proper folder structure. Without it, files will be stored with the prefix as part of the filename rather than in a folder.

  • ✅ Correct: tmp/ → Files stored as tmp/file.json
  • ❌ Incorrect: tmp → Files stored as tmpfile.json

Custom prefix examples:

# Production environment
export AWS_S3_TMP_PREFIX=prod/tmp/

# Staging environment
export AWS_S3_TMP_PREFIX=staging/tmp/

# Organize by date (requires manual updates)
export AWS_S3_TMP_PREFIX=tmp/2025/01/

# No prefix (store at bucket root - not recommended)
export AWS_S3_TMP_PREFIX=

AWS_S3_FILES_PREFIX

Purpose : S3 prefix (folder path) for Files API objects (OpenAI and Anthropic /v1/files endpoints)

Default : files/

Best Practice : Configure an AbortIncompleteMultipartUpload S3 lifecycle rule on this prefix to clean up abandoned upload parts, and apply Intelligent-Tiering for cost optimisation

export AWS_S3_FILES_PREFIX=files/

S3 Prefix Format

Prefix semantics (folder-style paths, trailing-slash requirement) are explained under AWS_S3_TMP_PREFIX and apply here identically.

Custom prefix examples:

# Production environment
export AWS_S3_FILES_PREFIX=prod/files/

# Staging environment
export AWS_S3_FILES_PREFIX=staging/files/

# No prefix (store at bucket root - not recommended)
export AWS_S3_FILES_PREFIX=

AWS_S3_VIDEOS_PREFIX

Purpose : S3 prefix (folder path) for videos generated through the Videos API

Default : videos/

Requirement : Must be non-empty, use only S3-safe characters (alphanumerics plus ! _ . * ' ( ) - per path segment), and end with a trailing / — an empty value would widen the ownership check that scopes listing/retrieval to the whole bucket

Best Practice : Generated videos persist until deleted through the API — configure an S3 lifecycle rule on this prefix to cap storage costs

export AWS_S3_VIDEOS_PREFIX=videos/

Amazon Bedrock writes each video generation job's output (MP4 and manifest) under this prefix, in a folder named after the job. Because Amazon Bedrock requires the output bucket to be in the same region as the invocation, videos are stored in the AWS_S3_REGIONAL_BUCKETS bucket of the region that served the job.

AWS_S3_VIDEOS_EXPIRES_AFTER

Purpose : Retention period in seconds (minimum 3600) for videos generated through the Videos API

Default : Unset — videos never expire and persist until deleted through the API

Best Practice : Pair with an S3 Lifecycle expiration rule on AWS_S3_VIDEOS_PREFIX covering the same duration (rounded up to whole days) so the objects are actually deleted

# Expire generated videos after 24 hours
export AWS_S3_VIDEOS_EXPIRES_AFTER=86400

When set, the Video object reports expires_at (job completion time plus this value) and downloading expired video content returns a 404. The server enforces expiry at the API level only; the paired S3 Lifecycle rule performs the physical cleanup.

AWS_S3_BATCHES_PREFIX

Purpose : S3 prefix (folder path) for the data of the Batch API and the Message Batches API — the submitted requests, the results, and the batch records themselves

Default : batches/

Requirement : Must be non-empty, use only S3-safe characters (alphanumerics plus ! _ . * ' ( ) - per path segment), and end with a trailing / — batches are addressed by prefix, so an empty value would make every stray object at the bucket root a candidate

Best Practice : Batch data persists until the batch is deleted — configure an S3 lifecycle rule on this prefix to cap storage costs, and grant AWS_BEDROCK_BATCH_ROLE_ARN read and write access under it

export AWS_S3_BATCHES_PREFIX=batches/

Each batch stores its data under a folder of its own below this prefix. Because the output bucket must be in the same region as the model that serves the batch, that data is stored in the AWS_S3_REGIONAL_BUCKETS bucket of the region that served it.

Vector Stores

The Vector Stores API needs an Amazon S3 vector bucket — a resource type of its own, created separately from a general purpose bucket — plus the general purpose bucket in AWS_S3_BUCKET, which holds the stores' records. The vector store endpoints answer 503 until both are configured, or until AWS_BEDROCK_KNOWLEDGE_BASE_IDS names a knowledge base to serve instead.

AWS_S3_VECTORS_BUCKET

Purpose : Name of the Amazon S3 vector bucket that holds the indexed content of every vector store

Default : None — the Vector Stores API is disabled

Requirement : Requires AWS_S3_BUCKET, which holds the vector store records; startup fails if only the vector bucket is set. The gateway's role needs the Vector Stores permissions on it

export AWS_S3_VECTORS_BUCKET=my-llm-vectors-us-east-1

Create the bucket yourself, then let the gateway create and delete the indexes inside it — one per vector store, removed when the store is deleted or expires.

AWS_S3_VECTORS_REGION

Purpose : AWS region holding AWS_S3_VECTORS_BUCKET

Default : The first AWS_BEDROCK_REGIONS entry

Requirement : Must be the bucket's own region. A vector bucket is a regional resource whose content is only reachable there, so this setting has no failover: if the region is unreachable, so is the Vector Stores API

export AWS_S3_VECTORS_REGION=us-east-1

AWS_S3_VECTOR_STORES_PREFIX

Purpose : S3 prefix (folder path) in AWS_S3_BUCKET for the Vector Stores API records — the stores, their attached files and their file batches

Default : vector_stores/

Best Practice : Keep it distinct from the other prefixes so a lifecycle rule written for one never reaches the records of another; these records are the stores themselves, so no expiration rule belongs on this prefix

export AWS_S3_VECTOR_STORES_PREFIX=vector_stores/

VECTOR_STORE_EMBEDDING_MODEL

Purpose : Model that turns the indexed files and the search queries into vectors

Default : amazon.titan-embed-text-v2:0

Requirement : Must be an embedding model available in your configured regions — see Embeddings API

export VECTOR_STORE_EMBEDDING_MODEL=amazon.titan-embed-text-v2:0

Each vector store records the model it was created with and keeps using it, so changing this setting only affects stores created afterwards. Existing stores keep answering exactly as before.

VECTOR_STORE_CHUNK_SIZE_TOKENS

Purpose : Default chunk size, in tokens, for files indexed without an explicit chunking_strategy

Default : 800 — the same default the upstream API applies

Requirement : Between 100 and 4096

export VECTOR_STORE_CHUNK_SIZE_TOKENS=800

A request's own chunking_strategy always wins, and a store created with one applies it to every file later attached without one. Chunk sizes are approximate — see Chunking.

VECTOR_STORE_CHUNK_OVERLAP_TOKENS

Purpose : Default number of tokens shared between consecutive chunks, for files indexed without an explicit chunking_strategy

Default : 400 — the same default the upstream API applies

Requirement : Must not exceed half of VECTOR_STORE_CHUNK_SIZE_TOKENS; startup fails otherwise

export VECTOR_STORE_CHUNK_OVERLAP_TOKENS=400

More overlap keeps a sentence split across two chunks findable from either one, at the cost of more chunks to embed and store.

AWS_SQS_VECTOR_STORE_QUEUE_URL

Purpose : URL of the Amazon SQS standard queue that carries vector store indexing work, so a file keeps being indexed when the server that accepted it stops

Default : None — indexing runs in the server that accepted the request, and a file being indexed when that server stops is reported as failed

Requirement : Requires AWS_S3_VECTORS_BUCKET. Must be a standard queue (not FIFO), with a dead-letter queue, and the gateway's role needs the Durable vector store indexing permissions on it

export AWS_SQS_VECTOR_STORE_QUEUE_URL=https://sqs.us-east-1.amazonaws.com/123456789012/stdapi-ai-indexing

Attaching a file records the work on the queue before answering, and every server reads from that queue: whichever one is still running finishes it. The region comes from the URL, so there is nothing else to configure. Create the queue yourself, in the same account.

The queue only ever carries identifiers — which store, which files, which batch — never file content, so no indexed data is stored a second time.

See Durable indexing for what a client observes, and Resilience for how it behaves during a deployment.

Give the queue a dead-letter queue

A file the gateway cannot index is retried a few times and then reported as failed. Without a dead-letter queue its message is dropped at that point; with one it is kept, so you can see what was refused. The gateway reads the queue's own redrive policy at startup and reports a queue that has none as a startup warning.

AWS_BEDROCK_KNOWLEDGE_BASE_IDS

Purpose : Comma-separated allowlist of Amazon Bedrock knowledge bases served through the Vector Stores API. Each allowlisted knowledge base is addressed as the vector store vs_kb_<knowledgeBaseId> on every /v1/vector_stores endpoint, and is returned by GET /v1/vector_stores next to the stores the server owns

Default : Empty — no knowledge base is addressable, and a vs_kb_... identifier is answered exactly as an unknown vector store is, so the allowlist cannot be probed

Requirement : Each knowledge base must already exist, in the first AWS_BEDROCK_REGIONS entry, be a Bedrock managed (MANAGED) or customer-managed (VECTOR) knowledge base, and the gateway's role needs the Knowledge Base Vector Stores permissions on it

export AWS_BEDROCK_KNOWLEDGE_BASE_IDS=ABCDE12345,FGHIJ67890/KLMNO13579

Write each entry as <knowledgeBaseId>, or as <knowledgeBaseId>/<dataSourceId> when the knowledge base has more than one data source; with a single data source the server resolves it itself.

Both kinds of document knowledge base are served, and a store behaves the same on either. A knowledge base connected to a structured data store (SQL) or backed by an Amazon Kendra GenAI index (KENDRA) is not a vector store and must not be allowlisted.

The knowledge base itself always stays yours: the server never creates one and never deletes one. It searches it, and manages the documents of its data source. Name, description, creation time and status are read from the knowledge base, and the requests that would change them are refused — see Knowledge Base Stores.

This setting is independent of AWS_S3_VECTORS_BUCKET: a deployment that sets only this one serves its allowlisted knowledge bases and creates no store of its own.

What a knowledge base costs

A knowledge base search costs more than a search on a store the server owns, and a knowledge base backed by an always-on vector database bills whether it is queried or not. Both backends are offered so the choice is yours — see Cost Management.

AWS_TRANSCRIBE_S3_BUCKET

Purpose : Temporary S3 bucket for transcription workflows

Default : Falls back to AWS_S3_BUCKET if not specified

Requirement : Must be in the same region as AWS_TRANSCRIBE_REGION when that is set; with the default multi-region behavior it serves the primary Bedrock region, and AWS_S3_REGIONAL_BUCKETS entries serve the other candidate regions

# If AWS_TRANSCRIBE_REGION is us-east-1
export AWS_TRANSCRIBE_S3_BUCKET=my-transcribe-temp-us-east-1

# If AWS_TRANSCRIBE_REGION is eu-west-1
export AWS_TRANSCRIBE_S3_BUCKET=my-transcribe-temp-eu-west-1

AWS_TRANSCRIBE_STREAM_LANGUAGES

Purpose : Languages a streamed transcription (stream=true) picks between when the request names none

Default : Empty — a request naming no language is transcribed once the whole recording has been read, and its language detected

Format : JSON array of two or more language codes; a single entry has no effect

A streamed transcription returns text before the recording has been fully read, which requires knowing which languages to expect. A request that names its language — or two or more expected languages — always gets one. Set this to extend the same behavior to requests that name neither, listing the languages your callers actually send. Listing more than five is not recommended, and two variants of the same language (en-US and en-GB) cannot both appear.

export AWS_TRANSCRIBE_STREAM_LANGUAGES='["en-US", "es-US", "fr-FR"]'

AWS_TRANSCRIBE_OUTPUT_ENCRYPTION_KEY_ARN

Purpose : AWS KMS key encrypting the transcription output written to AWS_TRANSCRIBE_S3_BUCKET

Default : None — output objects keep the bucket's own default encryption (SSE-S3)

Format : KMS key ARN: arn:<partition>:kms:<region>:<account-id>:key/<key-id>; startup fails on any other value

Requirement : The server's role needs kms:GenerateDataKey and kms:Decrypt on the key (IAM Permissions). With the default multi-region behavior the key must be usable from every candidate region — a multi-Region key or a single AWS_TRANSCRIBE_REGION

export AWS_TRANSCRIBE_OUTPUT_ENCRYPTION_KEY_ARN=arn:aws:kms:us-east-1:123456789012:key/12345678-1234-1234-1234-123456789012

Each job also sends its request identifiers (stdapi-ai.request_id, stdapi-ai.server_id, and stdapi-ai.user_id when the user is known) as the KMS encryption context, so a key policy can be conditioned on them.

AWS_S3_REGIONAL_BUCKETS

Purpose : Region-specific S3 buckets for Bedrock async and batch inference operations, and for staging attachments too large to travel inside a request

Default : Empty (no regional buckets configured)

Format : JSON object with region names as keys and bucket names as values

Requirement : Some Bedrock models require S3 buckets in the same region for async and batch inference operations

export AWS_S3_REGIONAL_BUCKETS='{"us-east-1": "my-bedrock-temp-us-east-1", "eu-west-1": "my-bedrock-temp-eu-west-1"}'

When to Use

Configure this setting when:

  • Using Bedrock async inference API
  • Using Bedrock batch inference API
  • Working with models that require regional S3 storage
  • Accepting chat, messages or responses requests with large attachments, or embedding requests with large inputs — those are staged in the bucket of the region serving the request, which then serves that request alone without failing over

If not specified for a region where async/batch operations are attempted, those operations may fail. Requests carrying an attachment larger than the model reads inline are refused with 413 when no region able to serve the model has a bucket.

Automatic Fallback

For the first region in AWS_BEDROCK_REGIONS (your primary region), if no regional bucket is specified, the service automatically falls back to AWS_S3_BUCKET. You only need to configure regional buckets for additional regions beyond your primary one.

Terraform Module

When using the Terraform module, regional S3 buckets are created automatically for each region in aws_bedrock_regions. The bucket names are exposed via the aws_s3_regional_buckets output and passed to the container as AWS_S3_REGIONAL_BUCKETS. No manual configuration required.

Best Practice

Apply the same S3 Bucket Lifecycle Configuration to these regional buckets as you would for the primary bucket to automatically clean up temporary files.

AWS_S3_ACCEPTED_BUCKETS

Purpose : Declare external S3 buckets that the application has read access to, mapped to their AWS region

Type : JSON object (keys: bucket names, values: AWS region identifiers)

Default : {} (empty — only the application's own buckets are recognized)

Behavior : Declaring a bucket here enables input access to objects the application does not own:

- **S3 URI and S3 HTTP URL access** — `s3://` URIs and S3 HTTP URLs (including presigned URLs) pointing at these buckets are accepted as input sources; an HTTP URL is automatically converted to an `s3://` URI so Bedrock can access the object directly.
- **Declared region** — The region mapped to each bucket is used to reach that bucket in its own region when reading the input object. It does not influence model or inference region selection.

Without this setting, only the application's own buckets (`AWS_S3_BUCKET` and `AWS_S3_REGIONAL_BUCKETS`) are recognized.
export AWS_S3_ACCEPTED_BUCKETS='{"my-data-bucket": "us-east-1", "my-eu-bucket": "eu-west-1"}'

Required IAM Permissions

The application's IAM role must have s3:GetObject permission on each declared bucket. Granting access at the bucket level is recommended:

{
  "Effect": "Allow",
  "Action": "s3:GetObject",
  "Resource": [
    "arn:aws:s3:::my-data-bucket/*",
    "arn:aws:s3:::my-eu-bucket/*"
  ]
}

When to Use

Configure this when your users provide S3 URLs from buckets outside the application's own buckets. This enables automatic HTTP-to-S3 URI conversion and optimal region routing for those objects.

S3 Bucket Lifecycle Configuration

Purpose : Configure automatic deletion of temporary files and abandoned multipart upload parts to minimize storage costs

Recommendation : Configure S3 lifecycle policies to automatically delete objects under the AWS_S3_TMP_PREFIX after 1 day, and abort incomplete multipart uploads under the AWS_S3_FILES_PREFIX after 1 day

stdapi.ai stores temporary files under the prefix configured by AWS_S3_TMP_PREFIX (default: tmp/). These include generated images, audio files, and transcription workflow files. Configure S3 lifecycle policies to automatically delete objects under this prefix after 1 day.

Additionally, multipart file uploads (OpenAI Uploads API) store parts under AWS_S3_FILES_PREFIX (default: files/). If a session is never completed or cancelled — for example when a client disconnects — the uploaded parts remain in S3 and accumulate costs. Add an AbortIncompleteMultipartUpload rule on the files prefix to clean these up automatically.

Application Cleanup Behavior

Short-lived temporary files: The application attempts to clean up short-lived temporary files (such as intermediate transcription files) after processing completes.

Results shared with clients: Files shared with clients using presigned URLs (such as generated images and audio) are never cleaned up automatically by the application. These files remain in S3 until removed by lifecycle policies or manual deletion.

Why lifecycle policies are essential: Since the application cannot determine when a client has finished using a presigned URL, S3 lifecycle policies are the recommended mechanism to clean up these files and prevent unbounded storage growth.

{
  "Rules": [
    {
      "Id": "DeleteTemporaryFiles",
      "Status": "Enabled",
      "Filter": {
        "Prefix": "tmp/"
      },
      "Expiration": {
        "Days": 1
      },
      "AbortIncompleteMultipartUpload": {
        "DaysAfterInitiation": 1
      }
    },
    {
      "Id": "AbortIncompleteMultipartUploads",
      "Status": "Enabled",
      "Filter": {
        "Prefix": "files/"
      },
      "AbortIncompleteMultipartUpload": {
        "DaysAfterInitiation": 1
      }
    }
  ]
}

Important: Update the Prefixes

The "Prefix" values in the lifecycle policy must match your AWS_S3_TMP_PREFIX and AWS_S3_FILES_PREFIX settings. If you use custom prefixes, update the policy accordingly.

Examples:

  • If AWS_S3_TMP_PREFIX=temporary/, use "Prefix": "temporary/" in the first rule
  • If AWS_S3_FILES_PREFIX=prod/files/, use "Prefix": "prod/files/" in the second rule

Apply via AWS CLI:

# For primary S3 bucket (AWS_S3_BUCKET)
aws s3api put-bucket-lifecycle-configuration \
  --bucket my-stdapi-bucket \
  --lifecycle-configuration file://lifecycle-policy.json

# For transcribe S3 bucket (AWS_TRANSCRIBE_S3_BUCKET, if different from AWS_S3_BUCKET)
aws s3api put-bucket-lifecycle-configuration \
  --bucket my-transcribe-temp-bucket \
  --lifecycle-configuration file://lifecycle-policy.json

# For regional buckets (AWS_S3_REGIONAL_BUCKETS)
aws s3api put-bucket-lifecycle-configuration \
  --bucket my-stdapi-us-west-2-bucket \
  --lifecycle-configuration file://lifecycle-policy.json

Apply to All S3 Buckets

Apply this lifecycle policy to:

  • AWS_S3_BUCKET - Primary bucket for generated files
  • AWS_TRANSCRIBE_S3_BUCKET - Transcription temporary files (if different from AWS_S3_BUCKET)
  • AWS_S3_REGIONAL_BUCKETS - All regional buckets for async/batch operations

All these buckets use the same AWS_S3_TMP_PREFIX for temporary file storage, and the same AWS_S3_FILES_PREFIX for multipart upload parts.

Bedrock Configuration

AWS_BEDROCK_REGIONS

Purpose : List of AWS regions where Bedrock models are available

Format : Comma-separated string

Default : Current AWS SDK region if not specified

Behavior : Models are discovered in the same order as the listed regions. The first region is the primary region where your server should be hosted on AWS for optimal performance. Your S3 bucket (AWS_S3_BUCKET) must also be in this region. If a model is unavailable in the primary region, subsequent regions are checked in order

export AWS_BEDROCK_REGIONS=us-east-1,us-west-2,eu-west-1

Region Selection Guide

Region Description
us-east-1 Widest model selection, usually gets latest releases first
us-west-2 Good selection, often early access to new models
eu-west-1 European compliance, subset of US models available

Advanced Configuration

See Compliance and Latency Optimization for detailed configuration examples including GDPR compliance, regional optimization strategies, and best practices for multi-region deployments.

Startup Warning

If any models in the configured regions fail availability checks (not enabled, unauthorized, or missing entitlement/agreement in your AWS account), a warning listing the affected models and per-region issues is logged at startup. Enable the required models in the Amazon Bedrock console for each configured region.

Unreachable Region Tolerance

A configured region that cannot be reached (invalid region for the account, network issue, throttling) does not block startup: it is skipped with an unreachable_bedrock_regions warning and its models are served from the remaining regions. The skipped region is retried automatically on the next model list refresh (see MODEL_CACHE_SECONDS), so a recovered region rejoins without a restart. Startup only fails when every configured region fails, or when every per-model availability check errors (e.g. the bedrock:GetFoundationModelAvailability permission is denied) — which indicates broken credentials or configuration rather than a regional outage.

AWS_BEDROCK_CROSS_REGION_INFERENCE

Purpose : Enable automatic cross-region routing when a model isn't available in the primary region

Type : Boolean

Default : true

export AWS_BEDROCK_CROSS_REGION_INFERENCE=true

AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL

Purpose : Allow global cross-region inference routing to any region worldwide

Type : Boolean

Default : true

GDPR Compliance

Set to false to comply with data residency regulations (e.g., EU GDPR) by restricting to regional inference only

export AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL=false

AWS_BEDROCK_REGION_ROUTING

Purpose : Automatic region routing strategy for distributing Bedrock requests across configured regions

Type : String

Default : ordered

Behavior : When multiple regions are configured in AWS_BEDROCK_REGIONS, this setting controls how requests are distributed across them. The router automatically handles quota/throttling errors and regional unavailability by temporarily avoiding affected regions

Requirement : Requires at least 2 regions in AWS_BEDROCK_REGIONS to take effect

Available strategies:

Strategy Description
disabled No routing; uses the single region where the model was discovered
ordered Try regions in configured order, skipping temporarily blocked ones (default). Best for prompt caching compatibility
lowest_latency Prefer the region with lowest measured latency. Latencies are measured at startup
round_robin Distribute requests evenly across regions. Incompatible with prompt caching
# Use ordered routing (default)
export AWS_BEDROCK_REGION_ROUTING=ordered

# Use lowest latency routing
export AWS_BEDROCK_REGION_ROUTING=lowest_latency

# Disable routing
export AWS_BEDROCK_REGION_ROUTING=disabled

Strategy Selection

  • ordered (default): Best general-purpose choice. Compatible with prompt caching since requests consistently go to the same region. Provides failover when a region hits quota limits
  • lowest_latency: Best when response time is critical. Measures region latencies at startup and prefers the fastest region. Falls back to others when the preferred region is blocked
  • round_robin: Best for maximizing aggregate throughput across regions. Not recommended with prompt caching as it distributes requests across all regions equally

More Details

For comprehensive documentation on region routing including failover behavior, S3 bucket pinning, logging, and best practices, see the Region Routing Guide.

AWS_BEDROCK_REGION_ROUTING_QUOTA_BACKOFF_SECONDS

Purpose : Duration to temporarily avoid a region after receiving a quota or throttling error

Type : Integer (seconds, must be > 0)

Default : 60

Behavior : This is the base backoff value. When a Bedrock API call fails due to quota limits (ThrottlingException, TooManyRequestsException, ServiceQuotaExceededException), the affected region is temporarily blocked. The actual delay doubles with each consecutive quota error on the same region (exponential backoff), up to a hard ceiling of 1 hour. The counter resets after a successful request. Subsequent requests are routed to other available regions during the backoff period.

# Default: 60 seconds (base value — actual delay doubles per consecutive error)
export AWS_BEDROCK_REGION_ROUTING_QUOTA_BACKOFF_SECONDS=60

# Shorter base backoff for aggressive retry
export AWS_BEDROCK_REGION_ROUTING_QUOTA_BACKOFF_SECONDS=30

# Longer base backoff for conservative approach
export AWS_BEDROCK_REGION_ROUTING_QUOTA_BACKOFF_SECONDS=120

Tuning

The base value controls how long the first quota error blocks a region. Subsequent consecutive errors on the same region double the delay (60 s → 120 s → 240 s → …, capped at 1 hour). Lower base values retry the region sooner but risk repeated throttling. Higher values provide more conservative avoidance at the cost of reduced region utilization.

See Region Routing — Overview for full backoff behavior details.

AWS_BEDROCK_REGION_ROUTING_UNAVAILABLE_BACKOFF_SECONDS

Purpose : Duration to temporarily avoid a region after receiving an unavailability error

Type : Integer (seconds, must be > 0)

Default : 30

Behavior : When a Bedrock API call fails due to service unavailability (ServiceUnavailableException, ModelNotReadyException), the affected region is temporarily blocked for this many seconds. These errors are typically shorter-lived than quota limits, so the default is shorter

# Default: 30 seconds
export AWS_BEDROCK_REGION_ROUTING_UNAVAILABLE_BACKOFF_SECONDS=30

# Longer backoff for stability
export AWS_BEDROCK_REGION_ROUTING_UNAVAILABLE_BACKOFF_SECONDS=60

More Details

See Region Routing — Overview for full backoff behavior details.

AWS_BEDROCK_REGION_ROUTING_MAX_QUOTA_BACKOFF_SECONDS

Purpose : Hard ceiling in seconds on the exponential quota backoff for a single region

Type : Integer (seconds, must be > 0)

Default : 3600 (1 hour)

Behavior : Quota backoff grows exponentially with consecutive errors (base interval × 2^n). This setting caps how large that value can become, preventing a region from being blocked indefinitely. Reduce it to allow faster recovery; increase it to keep a misbehaving region sidelined for longer.

# Default: 1 hour ceiling
export AWS_BEDROCK_REGION_ROUTING_MAX_QUOTA_BACKOFF_SECONDS=3600

# More aggressive recovery
export AWS_BEDROCK_REGION_ROUTING_MAX_QUOTA_BACKOFF_SECONDS=600

AWS_BEDROCK_REGION_ROUTING_QUOTA_STALE_FACTOR

Purpose : Multiplier applied to the max quota backoff to compute the stale-error reset threshold

Type : Integer (must be > 0)

Default : 2 (threshold = 2 × max quota backoff = 2 hours with defaults)

Behavior : If the most recent quota error on a region occurred more than max_quota_backoff × factor seconds ago, the consecutive-error counter is reset and the next error is treated as a fresh start rather than an escalation. A higher value keeps memory of past errors for longer before resetting the counter.

# Default: reset counter after 2× the max backoff window
export AWS_BEDROCK_REGION_ROUTING_QUOTA_STALE_FACTOR=2

# Longer memory of past errors
export AWS_BEDROCK_REGION_ROUTING_QUOTA_STALE_FACTOR=4

AWS_BEDROCK_MAX_RETRIES

Purpose : Maximum number of retries per Bedrock invocation, each retry escalating to the next available region

Type : Integer (must be 0 or greater; 0 disables retries)

Default : 9

Behavior : Controls the retry budget for each Bedrock API call. When region routing is enabled, every retry escalates to the next region in priority order and each candidate region is tried at most once, so the attempts are bounded by the smaller of AWS_BEDROCK_MAX_RETRIES + 1 and the number of candidate regions for the model — with 3 regions and the default 9 retries, a request makes at most 3 attempts. A region that just failed is still blocked by its own backoff, and retrying it would only extend that backoff instead of recovering the request. When routing is disabled, or the region is pinned by S3 inputs, the full budget is spent as SDK retries against that single region.

# Default: 9 retries (10 total attempts)
export AWS_BEDROCK_MAX_RETRIES=9

# Fail faster (e.g. low-latency interactive use cases)
export AWS_BEDROCK_MAX_RETRIES=3

# Deeper in-region retrying for single-region or S3-pinned requests
export AWS_BEDROCK_MAX_RETRIES=18

Related setting

See AWS_BEDROCK_REGION_ROUTING and Region Routing for the full retry and failover behavior.

AWS_FAILOVER_MAX_RETRIES

Purpose : Maximum SDK retry attempts per candidate region for the multi-region failover services (Polly, Transcribe, Translate, Comprehend)

Type : Integer (must be 0 or greater)

Default : 2

Behavior : Only applied when a service has several candidate regions (no explicit region setting): each region attempt uses this reduced retry budget (2 retries = 3 attempts per region) before failing over, so failover across regions replaces deep in-region retrying. When a service is pinned to a single region, the standard retry budget from AWS_BEDROCK_MAX_RETRIES applies instead.

# Default: 2 retries (3 attempts) per candidate region
export AWS_FAILOVER_MAX_RETRIES=2

# Fail over after a single attempt per region
export AWS_FAILOVER_MAX_RETRIES=0

Related setting

See Other AWS Services Failover for the full multi-region failover behavior.

AWS_BEDROCK_MANTLE_ENABLED

Purpose : Expose models served by the Amazon Bedrock Mantle endpoint (OpenAI/Anthropic-compatible APIs) in addition to the classic Bedrock Converse models

Type : Boolean

Default : true

Behavior : Mantle-only models (e.g. OpenAI GPT, xAI Grok, Google Gemma 4) become available on the chat completions, responses, messages, and completions routes. Models available on both the classic bedrock-runtime endpoint and Mantle are served by bedrock-runtime unless listed in AWS_BEDROCK_MANTLE_PREFERRED_MODELS.

Authentication requires no static secrets: short-term bearer tokens are derived automatically (SigV4-presigned) from the same AWS credential chain the server already uses, and refreshed transparently.

When Bedrock Mantle is unreachable or the IAM role lacks `bedrock-mantle` permissions, Mantle models are simply not listed and a warning is logged at startup — no configuration change required.
export AWS_BEDROCK_MANTLE_ENABLED=false

Guardrails Not Supported

Amazon Bedrock Guardrails are not supported on Mantle-served requests. When guardrails are configured while Mantle models are exposed, a startup warning reports how many models are affected; set AWS_BEDROCK_MANTLE_ENABLED=false to disable them.

Cross-Region Inference Profiles Not Available

Bedrock cross-region inference profiles do not exist on the Mantle endpoint. Mantle relies on multi-region failover and its own separate throughput quotas instead.

Required IAM Permissions

Enabling this setting requires the bedrock-mantle IAM permissions — see Bedrock Mantle IAM Permissions.

Bedrock Mantle Models feature overview

AWS_BEDROCK_MANTLE_REGIONS

Purpose : List of AWS regions used for Amazon Bedrock Mantle, in failover priority order

Type : Comma-separated string of AWS region identifiers

Default : The regions of AWS_BEDROCK_REGIONS that offer Bedrock Mantle

Behavior : Model availability differs per region; the served model catalog is the union of all listed regions. Region failover, quota backoff, and health tracking work exactly like classic Bedrock region routing.

export AWS_BEDROCK_MANTLE_REGIONS=us-east-1,eu-west-1

Regions Without a Mantle Endpoint

Bedrock Mantle is offered in fewer regions than classic Bedrock — see model availability by endpoint. Left unset, this setting keeps only the regions of AWS_BEDROCK_REGIONS known to offer it, so a deployment spanning other regions is not held up at startup by an address that does not exist.

An explicit value is used exactly as given, which is how a region AWS adds later is used without waiting for a release. A region that turns out to have no Mantle endpoint is named in a startup warning rather than retried forever. If none of your regions offers it, set AWS_BEDROCK_MANTLE_ENABLED=false.

AWS_BEDROCK_MANTLE_ENDPOINT_URL

Purpose : Override the Amazon Bedrock Mantle endpoint URL template

Type : String — URL template with a {region} placeholder

Default : None (https://bedrock-mantle.{region}.api.aws)

Behavior : The {region} placeholder is substituted with the target region.

export AWS_BEDROCK_MANTLE_ENDPOINT_URL='https://bedrock-mantle.{region}.api.aws'

AWS_BEDROCK_MANTLE_PREFERRED_MODELS

Purpose : Model IDs (or ID prefixes) served by Amazon Bedrock Mantle even when also available on the classic bedrock-runtime endpoint

Type : Comma-separated string of model IDs or ID prefixes

Default : [] (empty — dual-homed models are served by bedrock-runtime)

Behavior : Useful to leverage Mantle's independent throughput quotas or native response storage for selected models. Mantle quotas (per-model, per-region tokens-per-minute) are independent from bedrock-runtime quotas.

export AWS_BEDROCK_MANTLE_PREFERRED_MODELS='anthropic.claude-haiku-4-5,openai.gpt-oss'

AWS_BEDROCK_MANTLE_SERVICE_HEADER

Purpose : Honor the x-stdapi-service: bedrock-mantle request header to route a model available on both endpoints through Bedrock Mantle for that request instead of the default bedrock-runtime serving

Type : Boolean

Default : false

export AWS_BEDROCK_MANTLE_SERVICE_HEADER=true

Incompatible with Bedrock Guardrails

Requires AWS_BEDROCK_MANTLE_ENABLED and cannot be enabled together with Amazon Bedrock Guardrails: guardrails do not apply to Mantle-served requests, so a per-request header would allow clients to bypass them.

AWS_BEDROCK_MANTLE_PROJECT

Purpose : Default Amazon Bedrock Project/Workspace ID attributed to Bedrock Mantle inference requests for cost tracking and observability

Type : String — a bare project ID (e.g. proj_abc123 or default), not an ARN

Default : None (requests fall to the account's default project)

Behavior : Bedrock Projects (OpenAI-compatible APIs) and Workspaces (Anthropic Messages API) are the same underlying resource; the value is sent as the OpenAI-Project header on the Chat Completions and Responses APIs, and as the anthropic-workspace header on the Anthropic Messages API. When unset, requests fall to the account's default project — no failure.

export AWS_BEDROCK_MANTLE_PROJECT=proj_abc123

Bedrock Mantle only

Project/Workspace attribution is honored only for models served by the Amazon Bedrock Mantle endpoint. Classic bedrock-runtime (non-Mantle) models ignore it and use application inference profiles instead.

AWS_BEDROCK_ALLOW_MANTLE_PROJECT_OVERRIDE

Purpose : Allow a request to override the configured Mantle project via the OpenAI-Project / anthropic-workspace header

Type : Boolean

Default : false

Behavior : When true, a request may set its own project through the OpenAI-Project (Chat Completions, Responses) or anthropic-workspace (Anthropic Messages) header. When false and AWS_BEDROCK_MANTLE_PROJECT is configured, the request header is ignored and the server default applies. When no default project is configured, the request header is always honored regardless of this flag. A malformed request-supplied project ID returns 400.

export AWS_BEDROCK_ALLOW_MANTLE_PROJECT_OVERRIDE=true

Bedrock Mantle only

These headers apply only to models served by the Amazon Bedrock Mantle endpoint; classic bedrock-runtime models ignore them.

AWS_BEDROCK_EXTERNAL_WEB_ACCESS

Purpose : Let the built-in web search tool reach the public web

Type : Boolean

Default : false

Behavior : Controls whether the built-in web search tool may reach the external web. Searches are answered from the Amazon Bedrock web index and cache either way, and answers are current and carry source citations. AWS documents that retrieval is served entirely from that index and cache today, so no request data leaves the AWS boundary even when this is enabled, and that a future release may allow live external retrieval — at which point request data may leave it. Enabling it is therefore a decision taken in advance about behaviour that can change. It also requires the bedrock-websearch:ExternalWebAccess IAM permission on the credentials this server uses; each action is authorized only when a model actually attempts it, and a denied call does not fail the request.

export AWS_BEDROCK_EXTERNAL_WEB_ACCESS=true

AWS_BEDROCK_ALLOW_EXTERNAL_WEB_ACCESS_OVERRIDE

Purpose : Allow a request to override AWS_BEDROCK_EXTERNAL_WEB_ACCESS with the external_web_access extra model parameter

Type : Boolean

Default : false

Behavior : When true, a request that sends external_web_access as an extra model parameter decides for that request, on the models whose web search takes a web access choice per request — the OpenAI GPT-5.x family. When false, a request that sets it to anything other than the configured value is rejected with 400 rather than being silently overridden; a request that omits it always gets the configured value. A request asking for a value a model cannot be given is rejected with 400 as well, rather than accepted and quietly ignored. On the models that do take it, a request naming no web search tool is accepted: nothing is searched, so the value has no search to apply to.

export AWS_BEDROCK_ALLOW_EXTERNAL_WEB_ACCESS_OVERRIDE=true

AWS_BEDROCK_MODEL_REGION_RESTRICT

Purpose : Restrict a model to specific region(s) only, useful when a model provides important features only in certain regions

Type : JSON object (keys: Bedrock model IDs or prefixes, values: ordered lists of allowed regions)

Default : {} (empty — no model-specific region restriction)

Behavior : When set, the model is made available only in the listed regions (intersected with the regions where it is actually available), and the list order defines the routing priority when the default ordered routing strategy is used. No fallback to other regions occurs. Keys can be exact model IDs or prefixes that match the beginning of a model ID

# Restrict Nova Pro to us-east-1 for grounding support
export AWS_BEDROCK_MODEL_REGION_RESTRICT='{"amazon.nova-pro-v1:0": ["us-east-1"]}'

Use Case: Region-Specific Features

Some model features are only available in specific regions. For example, Nova grounding is only available in us-east-1. Restricting the model to that region ensures the feature is always available.

See Region Routing — Model Region Restrict for more details.

Startup Warning

If a key has no matching available model, a warning is logged at startup. This can happen for two reasons:

  • Typo or unknown model — the key (exact ID or prefix) does not match any model ID returned by Bedrock.
  • No matching region — the model exists but is not available in any of the regions listed in AWS_BEDROCK_REGIONS (e.g. the model is not enabled in those regions, or the restricted regions are not configured).

AWS_BEDROCK_LEGACY

Purpose : Allow usage of legacy/deprecated Bedrock models

Type : Boolean

Default : false

export AWS_BEDROCK_LEGACY=true

AWS_BEDROCK_DEPRECATED_MODEL_FALLBACK

Purpose : Transparently reroute requests using a deprecated model ID to its recommended replacement

Type : Boolean

Default : true

Behavior : When true, any request that specifies a deprecated model ID (as listed in the server's deprecation registry) is silently retried with the recommended replacement model. The replacement is fully re-evaluated — alias resolution, modality checks, and region routing all apply to the new model ID. When false, deprecated model IDs return a 404 error with a message indicating the replacement, forcing clients to migrate explicitly.

# Transparent fallback (default) — clients using old model IDs keep working
export AWS_BEDROCK_DEPRECATED_MODEL_FALLBACK=true

# Strict mode — deprecated model IDs return 404, clients must update their code
export AWS_BEDROCK_DEPRECATED_MODEL_FALLBACK=false

AWS_BEDROCK_DEPRECATED_MODELS

Purpose : Extend or override the built-in deprecated model registry with custom mappings

Type : JSON object — dict[str, str]

Default : {}

Behavior : Merged with the built-in registry at startup. User-provided entries take precedence over built-in ones — this means it can be used both to add new deprecated model mappings and to override the fallback target of an already-defined deprecated model. Effective only when AWS_BEDROCK_DEPRECATED_MODEL_FALLBACK is true.

Reference : Amazon Bedrock model lifecycle

# Add a custom deprecated model and override an existing built-in mapping
export AWS_BEDROCK_DEPRECATED_MODELS='{"my-old-model-v1": "my-new-model-v2", "amazon.titan-text-lite-v1": "amazon.nova-lite-v1:0"}'

AWS_BEDROCK_MARKETPLACE_AUTO_SUBSCRIBE

Purpose : Control automatic subscription to new models in AWS Marketplace

Type : Boolean

Default : true

Behavior : When true, the server automatically subscribes to new models discovered in the AWS Marketplace, making them immediately available through the API. When false, only models with existing marketplace subscriptions are visible and accessible

IAM Permissions Required : aws-marketplace:Subscribe, aws-marketplace:ViewSubscriptions — see Marketplace Auto-Subscribe IAM

# Allow automatic subscription (default)
export AWS_BEDROCK_MARKETPLACE_AUTO_SUBSCRIBE=true

# Restrict to pre-subscribed models only
export AWS_BEDROCK_MARKETPLACE_AUTO_SUBSCRIBE=false

What is Marketplace Auto-Subscribe?

Amazon Bedrock requires marketplace subscription before certain models can be used. This setting controls whether stdapi.ai automatically handles the subscription process:

  • true (default): Models are automatically subscribed when discovered, providing seamless access to new models as they become available
  • false: Only models that have already been subscribed through the AWS Marketplace are visible, providing explicit control over model access

When to Disable

Set to false when:

  • You need explicit control over which models are accessible
  • You want to prevent automatic marketplace subscriptions that may incur costs
  • Your organization requires manual approval for new AI model usage
  • Compliance policies require pre-authorization of AI models

AWS Documentation

For more information about Bedrock model access and marketplace registration, see the Amazon Bedrock Model Access documentation.

AWS_BEDROCK_ALLOW_CROSS_REGION_INFERENCE_PROFILE_ARN

Purpose : Allow users to pass cross-region inference profile ARNs directly as model IDs in API requests

Type : Boolean

Default : false

Behavior : When enabled, users can use cross-region inference profile ARNs instead of model IDs in the model parameter. Cross-region inference profiles enable routing to multiple regions for better availability

IAM Permissions Required : bedrock:GetInferenceProfile (see IAM Permissions)

# Disabled (default) - users can only use standard model IDs
# No environment variable needed

# Enable cross-region inference profile ARN support
export AWS_BEDROCK_ALLOW_CROSS_REGION_INFERENCE_PROFILE_ARN=true

Additional IAM Permissions Required

Enabling this setting requires adding the bedrock:GetInferenceProfile IAM permission to your role/user. Without this permission, API requests using inference profile ARNs will fail with authorization errors.

See the Bedrock Inference Profiles and Prompt Routers IAM section for the complete policy configuration.

Example ARN

arn:aws:bedrock:us-east-1:123456789012:inference-profile/us.anthropic.claude-sonnet-5

Automatic Cross-Region Routing (Default Behavior)

By default, stdapi.ai automatically determines and uses the best cross-region inference profile for each model, based on AWS_BEDROCK_REGIONS, AWS_BEDROCK_CROSS_REGION_INFERENCE, and AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL. Manually passing cross-region inference profile ARNs is only needed in rare cases to override that selection — for most deployments, leave this disabled. See Using Inference Profile and Prompt Router ARNs for details.

AWS_BEDROCK_ALLOW_APPLICATION_INFERENCE_PROFILE_ARN

Purpose : Allow users to pass application inference profile ARNs directly as model IDs in API requests

Type : Boolean

Default : false

Behavior : When enabled, users can use application inference profile ARNs instead of model IDs in the model parameter. Application inference profiles are custom routing configurations for specific use cases

IAM Permissions Required : bedrock:GetInferenceProfile (see IAM Permissions)

# Disabled (default) - users can only use standard model IDs
# No environment variable needed

# Enable application inference profile ARN support
export AWS_BEDROCK_ALLOW_APPLICATION_INFERENCE_PROFILE_ARN=true

Additional IAM Permissions Required

Enabling this setting requires adding the bedrock:GetInferenceProfile IAM permission to your role/user. Without this permission, API requests using application inference profile ARNs will fail with authorization errors.

See the Bedrock Inference Profiles and Prompt Routers IAM section for the complete policy configuration.

Example ARN

arn:aws:bedrock:us-east-1:123456789012:application-inference-profile/abc123xyz

What are Application Inference Profiles?

Application inference profiles are custom routing configurations that you create in your AWS account. They allow you to define specific routing behavior, region preferences, and failover strategies tailored to your application's needs.

When to Enable

Enable this setting when:

  • You have custom application inference profiles configured in your AWS account
  • You need application-specific routing configurations
  • You want to give users access to custom profiles you've created

AWS_BEDROCK_ALLOW_PROMPT_ROUTER_ARN

Purpose : Allow users to pass prompt router ARNs directly as model IDs in API requests

Type : Boolean

Default : false

Behavior : When enabled, users can use prompt router ARNs instead of model IDs in the model parameter. Prompt routers enable dynamic model selection based on prompt characteristics

IAM Permissions Required : bedrock:GetPromptRouter (see IAM Permissions)

# Disabled (default) - users can only use standard model IDs
# No environment variable needed

# Enable prompt router ARN support
export AWS_BEDROCK_ALLOW_PROMPT_ROUTER_ARN=true

Additional IAM Permissions Required

Enabling this setting requires adding the bedrock:GetPromptRouter IAM permission to your role/user. Without this permission, API requests using prompt router ARNs will fail with authorization errors.

See the Bedrock Inference Profiles and Prompt Routers IAM section for the complete policy configuration.

Example ARN

arn:aws:bedrock:us-east-1:123456789012:default-prompt-router/my-router

What are Prompt Routers?

Prompt routers are intelligent routing systems that analyze prompt characteristics (length, complexity, language) and dynamically select the most appropriate model. This enables cost optimization and performance tuning based on request patterns.

When to Enable

Enable this setting when:

  • You have prompt routers configured in your AWS account
  • You want intelligent cost optimization through dynamic model selection
  • You need automatic model selection based on prompt complexity

AWS_BEDROCK_ALLOW_PROMPT_ARN

Purpose : Allow users to reference an Amazon Bedrock Prompt Management prompt ARN in the OpenAI Responses API prompt parameter

Type : Boolean

Default : false

Behavior : When enabled, prompt.id accepts a prompt ARN (with an optional prompt.version) and prompt.variables fill in the template. Amazon Bedrock renders the stored prompt, and the model it is bound to serves the request. When disabled, any prompt parameter is rejected with a 400 error

IAM Permissions Required : bedrock:GetPrompt (resolve the prompt's model) and bedrock:RenderPrompt (invoke it)

# Disabled (default) - the Responses API `prompt` parameter returns 400
# No environment variable needed

# Enable Prompt Management prompt ARN support
export AWS_BEDROCK_ALLOW_PROMPT_ARN=true

Additional IAM Permissions Required

Enabling this setting requires adding the bedrock:GetPrompt and bedrock:RenderPrompt IAM permissions, scoped to the prompt resources you want to expose. Without them, requests using a prompt ARN fail with authorization errors.

Example ARN

arn:aws:bedrock:us-east-1:123456789012:prompt/ABCDE12345:1

Scope and Limitations

  • Only TEXT prompts are supported, and the request's model must be the model the prompt is bound to.
  • Prompt variable values must be plain strings.
  • input, instructions, tools, text, previous_response_id and inference parameters cannot be combined with prompt.

See Managed Prompt Templates for the full request contract.

AWS_BEDROCK_MODEL_ARN_MAPPING

Purpose : Map standard model IDs to custom inference profile or prompt router ARNs for server-controlled routing

Format : JSON object with model IDs as keys and ARNs as values

Default : {} (empty, no mappings)

Behavior : When configured, the mapped ARN is used instead of the default cross-region inference profile when clients request the model by its standard ID. This provides centralized control over model routing without requiring client changes

export AWS_BEDROCK_MODEL_ARN_MAPPING='{
  "anthropic.claude-sonnet-5": "arn:aws:bedrock:us-east-1:123456789012:application-inference-profile/my-custom-profile",
  "anthropic.claude-haiku-4-5-20251001-v1:0": "arn:aws:bedrock:us-east-1:123456789012:default-prompt-router/my-router"
}'

What is Model ARN Mapping?

Model ARN mapping allows server administrators to override the default routing behavior for specific models. When a client requests a model using its standard ID (e.g., anthropic.claude-sonnet-5), the server automatically uses the mapped ARN for routing instead.

Supported ARN Types:

  • Cross-region inference profiles - AWS-managed multi-region routing
  • Application inference profiles - Custom routing configurations
  • Prompt routers - Intelligent dynamic model selection

Key Benefits

  • Centralized Control - Change routing behavior without modifying client code
  • Transparent to Clients - Clients use standard model IDs, server handles routing
  • Easy Migration - Switch between routing strategies by updating server config
  • Environment-Specific - Different mappings for dev/staging/production environments

Use Cases

Cost Optimization with Prompt Router:

export AWS_BEDROCK_MODEL_ARN_MAPPING='{
  "anthropic.claude-sonnet-5": "arn:aws:bedrock:us-east-1:123456789012:default-prompt-router/cost-optimizer"
}'
Automatically route simple prompts to cheaper models, complex prompts to premium models.

Custom Application Profile:

export AWS_BEDROCK_MODEL_ARN_MAPPING='{
  "anthropic.claude-sonnet-5": "arn:aws:bedrock:us-east-1:123456789012:application-inference-profile/production-profile"
}'
Use your custom inference profile with specific region preferences and failover behavior.

Environment-Specific Routing:

# Production: Use cost-optimized prompt router
export AWS_BEDROCK_MODEL_ARN_MAPPING='{"anthropic.claude-sonnet-5": "arn:aws:bedrock:us-east-1:123456789012:default-prompt-router/prod-router"}'

# Development: Use standard cross-region profile
export AWS_BEDROCK_MODEL_ARN_MAPPING='{}'

Best Practices

  • Test mappings in development before deploying to production
  • Document your ARN mappings and their purposes
  • Keep ARN mappings in version control alongside other configuration
  • Monitor routing behavior after updating mappings

Startup Warning

If any model IDs in AWS_BEDROCK_MODEL_ARN_MAPPING are not found among available Bedrock models, a warning listing the affected entries is logged at startup. This typically means the model is not enabled in your configured regions or the model ID contains a typo.

Other AWS Services

Optional Configuration

Each service region is optional. Left unset, the service treats every AWS_BEDROCK_REGIONS entry as a candidate and fails over between them; setting one pins the service to that single region, with no failover.

AWS_POLLY_REGION

Purpose : Region for Amazon Polly text-to-speech service

Default : All regions in AWS_BEDROCK_REGIONS, with per-engine regional discovery and automatic failover

Behavior : When unset, voice availability is discovered per engine in every AWS_BEDROCK_REGIONS entry at startup: an engine (Standard, Neural, Long-form, Generative) is exposed as a model when at least one candidate region offers it, and each synthesis call routes to the regions offering the requested engine and voice, failing over on region-level errors. Setting an explicit region pins Polly to that single region — engines it does not offer are then disabled.

export AWS_POLLY_REGION=us-east-1

Amazon Polly Engine Availability

Not all Polly engines (Standard, Neural, Long-form, Generative) are available in all AWS regions. With the default multi-region behavior, an engine missing from one region is simply served from another candidate region that offers it. See Amazon Polly feature and region compatibility for detailed information.

AWS_COMPREHEND_REGION

Purpose : Region for the Amazon Comprehend services (language detection and toxicity moderation)

Default : All regions in AWS_BEDROCK_REGIONS, tried in order with automatic failover

Behavior : When unset, Comprehend calls try each AWS_BEDROCK_REGIONS entry in order and fail over to the next region on region-level errors (throttling, service unavailability, network issues, or a region that does not offer Comprehend or the requested operation). Setting an explicit region pins Comprehend to that single region with no failover.

export AWS_COMPREHEND_REGION=us-east-1

Amazon Comprehend Regional Availability

Amazon Comprehend is not available in all AWS regions. stdapi.ai uses the detect_dominant_language feature for language detection and detect_toxic_content for Comprehend moderation. Verify service and feature availability in your target region (with the default multi-region behavior, a region without Comprehend simply fails over to the next one). See Amazon Comprehend supported regions for regional availability.

AWS_TRANSCRIBE_REGION

Purpose : Region for Amazon Transcribe speech-to-text service

Default : All regions in AWS_BEDROCK_REGIONS that have a co-located S3 bucket, tried in order with automatic failover

Behavior : Transcription jobs need an S3 bucket in the job's region. When unset, every AWS_BEDROCK_REGIONS entry with a usable bucket is a candidate — the primary region is served by AWS_TRANSCRIBE_S3_BUCKET (or AWS_S3_BUCKET), the others by their AWS_S3_REGIONAL_BUCKETS entry. On a region-level error while starting a job, the audio is server-side copied to the next candidate's bucket and the job restarts there. Setting an explicit region pins Transcribe to that single region with no failover.

export AWS_TRANSCRIBE_REGION=us-east-1

AWS_TRANSLATE_REGION

Purpose : Region for Amazon Translate text translation service

Default : All regions in AWS_BEDROCK_REGIONS, tried in order with automatic failover

Behavior : When unset, translation calls try each AWS_BEDROCK_REGIONS entry in order and fail over to the next region on region-level errors (throttling, service unavailability, network issues, or a region that does not offer Translate). Setting an explicit region pins Translate to that single region with no failover.

export AWS_TRANSLATE_REGION=us-east-1

Compliance and Latency Optimization

Strategic region configuration is critical for both regulatory compliance and performance optimization. This section provides best practice configurations for common scenarios.

AWS AI Services Data Privacy

Amazon Bedrock: Does not store or use user prompts and responses, and does not share them with third parties by default. Your content remains private and is not used to train models.

Other AI Services: AWS collects telemetry data from other AI services (Polly, Comprehend, Transcribe, Translate) by default. For enhanced data privacy and compliance, you can opt out of AWS using your content to improve AI services. Configure AI services opt-out policies at the AWS Organizations level to prevent your data from being used for service improvement.

GDPR and Data Residency Compliance

For applications serving European users, data residency regulations like GDPR may require that data processing occurs within specific geographic boundaries.

EU-Only Configuration (Strict GDPR)
# Use only European regions
export AWS_S3_BUCKET=my-stdapi-eu-bucket
export AWS_BEDROCK_REGIONS=eu-west-1,eu-west-3,eu-central-1

# Disable global cross-region inference to prevent data routing outside Europe
export AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL=false

# Keep cross-region inference enabled for failover within EU regions
export AWS_BEDROCK_CROSS_REGION_INFERENCE=true

Key Compliance Settings

  • AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL=false: Prevents requests from being routed to regions outside your specified list
  • AWS_BEDROCK_CROSS_REGION_INFERENCE=true: Enables cross-region inference within your specified EU regions
  • All services in EU regions: Ensures all data processing stays within European boundaries

Important Considerations

  • Not all Bedrock models are available in all EU regions - verify model availability
  • Some newer models may be available in US regions first; this configuration prioritizes compliance over immediate access to latest models
  • S3 buckets must be created in EU regions and configured appropriately for data residency

Latency Optimization

For applications prioritizing low latency and high performance, configure regions closest to your users and application infrastructure.

🇺🇸 North America:

# Primary region for lowest latency, with fallbacks
export AWS_S3_BUCKET=my-stdapi-us-east-1-bucket
export AWS_BEDROCK_REGIONS=us-east-1,us-west-2,us-east-2

# Enable all cross-region inference for maximum model availability
export AWS_BEDROCK_CROSS_REGION_INFERENCE=true
export AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL=true

🇯🇵 Asia-Pacific:

# Use Asia-Pacific regions for lowest latency to APAC users
export AWS_S3_BUCKET=my-stdapi-ap-southeast-1-bucket
export AWS_BEDROCK_REGIONS=ap-southeast-1,ap-northeast-1,us-west-2

# Enable global inference for fallback to US regions if needed
export AWS_BEDROCK_CROSS_REGION_INFERENCE=true
export AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL=true

🌍 Global Multi-Region:

# Balanced configuration with worldwide coverage
export AWS_S3_BUCKET=my-stdapi-us-east-1-bucket
export AWS_BEDROCK_REGIONS=us-east-1,eu-west-1,ap-southeast-1,us-west-2

# Enable global inference for best availability
export AWS_BEDROCK_CROSS_REGION_INFERENCE=true
export AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL=true

Latency Optimization Tips

  • Server and S3 co-location: Deploy stdapi.ai and your AWS_S3_BUCKET in the first region specified in AWS_BEDROCK_REGIONS (your primary region)
  • Network proximity: Choose the first region based on low latency to your application servers and end users
  • Data transfer costs: Cross-region data transfer incurs costs; co-locating server and S3 in the same region minimizes these
  • Model availability: While us-east-1 often has the most models, check specific model availability in your target regions

Hybrid Approach: Compliance with Performance

Balance compliance requirements with performance needs:

EU Primary with US Fallback
# EU primary with US fallback (for model availability)
export AWS_S3_BUCKET=my-stdapi-eu-bucket
export AWS_BEDROCK_REGIONS=eu-west-1,eu-central-1,us-east-1

# Allow cross-region but restrict to specific regions only
export AWS_BEDROCK_CROSS_REGION_INFERENCE=true
export AWS_BEDROCK_CROSS_REGION_INFERENCE_GLOBAL=false

Legal Compliance Notice

Including us-east-1 as a fallback region provides access to more models but may not comply with strict data residency requirements. Consult your legal and compliance teams before using this configuration.


Configuration Order

When deploying stdapi.ai, configure settings in this recommended order:

  1. IAM Permissions - Set up AWS access first
  2. AWS Services and Regions - Configure S3 buckets and Bedrock regions
  3. Authentication - Secure your API with authentication
  4. Optional Features - Add observability, guardrails, and other features as needed

IAM Permissions

The full IAM reference — required Amazon Bedrock permissions, per-feature policy statements, complete policy examples, and AWS tag policy requirements — has moved to the dedicated IAM Permissions page.


Authentication

stdapi.ai supports three sources for API key authentication, plus Amazon Cognito user pool tokens.

API Key Sources

Configure exactly one source. If several are set, the first match in this precedence order is used and the others are ignored:

  1. Direct API keyAPI_KEY (highest precedence)
  2. SSM Parameter StoreAPI_KEY_SSM_PARAMETER
  3. Secrets ManagerAPI_KEY_SECRETSMANAGER_SECRET (lowest precedence)

The methods below are listed in that precedence order. SSM Parameter Store remains the recommended method for production.

Amazon Cognito user pool tokens (Method 4) are independent of the API key and can be accepted alongside it, or instead of it — AUTHENTICATION_MODE decides.

Conflicting Configuration

Only one combination is rejected at startup: API_KEY set together with a Secrets Manager source (API_KEY_SECRETSMANAGER_SECRET). Every other combination starts normally and is resolved silently by the precedence order above — the lower-precedence sources are never read.

No Authentication Warning

If neither an API key source nor a user pool is configured, the API accepts all requests without authentication and a security warning is logged at startup. This is suitable only for internal/private deployments.

Method 1: Direct API Key

Provide the API key directly via environment variable. Intended for local development and testing; it takes precedence over both AWS-backed sources.

API_KEY

Purpose : Static API key value

Security Warning : Avoid hardcoding in configuration files; use environment variables only

Client Usage : Clients must include this key in the Authorization: Bearer <key> header or X-API-Key header

export API_KEY=sk-1234567890abcdef...

Recommended - Use AWS Systems Manager Parameter Store for secure key storage with encryption, access control, and auditing. This method should be used only with already existing parameters.

API_KEY_SSM_PARAMETER

Purpose : Name of the SSM parameter containing the API key. The parameter is retrieved from the current region detected by the running container, or defaults to the first region in AWS_BEDROCK_REGIONS.

Recommendation : Use SecureString type for encryption at rest

IAM Permissions Required : ssm:GetParameter, kms:Decrypt (if encrypted)

export API_KEY_SSM_PARAMETER=/stdapi/prod/api-key

Method 3: Secrets Manager

Use AWS Secrets Manager for secure key storage with automatic rotation support. This method should be used only with already existing secrets.

API_KEY_SECRETSMANAGER_SECRET

Purpose : Name of the Secrets Manager secret containing the API key. The secret is retrieved from the current region detected by the running container, or defaults to the first region in AWS_BEDROCK_REGIONS.

Format : Can be a plain string or JSON object

IAM Permissions Required : secretsmanager:GetSecretValue

API_KEY_SECRETSMANAGER_KEY

Purpose : JSON key name within the secret (if the secret is a JSON object)

Default : api_key

Plain String Secret:

export API_KEY_SECRETSMANAGER_SECRET=stdapi-api-key

JSON Secret:

export API_KEY_SECRETSMANAGER_SECRET=stdapi-credentials
export API_KEY_SECRETSMANAGER_KEY=api_key

Example JSON secret structure:

{
  "api_key": "sk-1234567890abcdef...",
  "other_config": "value"
}

Method 4: Amazon Cognito User Pool Tokens

Accept the bearer tokens issued by an Amazon Cognito user pool instead of, or alongside, the API key. Each caller gets its own short-lived credential, and the verified caller is the identity per-user cost attribution bills against; withdrawing a caller's access takes effect when their current token expires. Clients send the token in the Authorization: Bearer <token> or X-API-Key header, like an API key. What is validated on every request is described in Authentication & Security.

export AWS_COGNITO_USER_POOL_ID=eu-west-3_a1b2c3d4e
export AWS_COGNITO_CLIENT_IDS=1example23456789abcdefghij

Incomplete configuration fails startup

A user pool without AWS_COGNITO_CLIENT_IDS, a Cognito setting without a pool, or an AUTHENTICATION_MODE that contradicts what is configured, all stop the server at startup with an explicit message — a partially configured pool never degrades into an unauthenticated deployment.

The pool also configures agent discovery

Add OAUTH_RESOURCE_IDENTIFIER and an AI agent can authenticate itself against the deployment. Nothing else is needed: the pool's issuer and required scopes are what get published — see Authentication Discovery for Agents.

AUTHENTICATION_MODE

Purpose : Which client authentication methods the deployment accepts

Default : any — every method that is configured

Values : - any: the API key and user pool tokens, whichever is configured - api_key: the API key only; startup fails if a user pool is also configured - cognito: user pool tokens only; startup fails if an API key source is also configured

export AUTHENTICATION_MODE=cognito

AWS_COGNITO_USER_POOL_ID

Purpose : Identifier of the user pool whose tokens authenticate clients. Setting it enables the method; the pool's AWS Region is read from the identifier itself, and the public signing keys are loaded from that Region at startup.

Default : None — user pool tokens are not accepted

Requirement : AWS_COGNITO_CLIENT_IDS must be set too

IAM Permissions Required : None — the signing keys are public

export AWS_COGNITO_USER_POOL_ID=eu-west-3_a1b2c3d4e

AWS_COGNITO_CLIENT_IDS

Purpose : Comma-separated app client IDs whose tokens are accepted. A token issued to any other app client of the pool is rejected.

Default : Empty — startup fails when a user pool is configured without it

Requirement : Required whenever AWS_COGNITO_USER_POOL_ID is set

export AWS_COGNITO_CLIENT_IDS=1example23456789abcdefghij,2example3456789abcdefghijk

AWS_COGNITO_REQUIRED_SCOPES

Purpose : Comma-separated OAuth 2.0 scopes a token must all carry to be accepted

Default : None — any scope set is accepted

Requirement : Custom scopes exist only on tokens issued by the pool's OAuth 2.0 token endpoint, which needs a resource server and a pool domain. Tokens obtained by signing in with a username and password carry only aws.cognito.signin.user.admin and are rejected when a custom scope is required.

export AWS_COGNITO_REQUIRED_SCOPES=stdapi/invoke

AWS_COGNITO_ACCEPT_ID_TOKEN

Purpose : Also accept identity tokens, not only access tokens

Default : false

Effect : Identity tokens describe the signed-in user rather than granting API access, and carry no scopes. Enable only for clients that cannot obtain an access token.

export AWS_COGNITO_ACCEPT_ID_TOKEN=true

AWS_COGNITO_ISSUER_TYPE

Purpose : The pool's issuer configuration, which decides the issuer URL its tokens carry

Default : original

Values : - original: https://cognito-idp.<region>.amazonaws.com/<pool-id> - updated: https://issuer-cognito-idp.<region>.amazonaws.com/<pool-id>, available on the Essentials and Plus pool tiers

Requirement : Must match the pool's own setting; tokens whose issuer differs are rejected

export AWS_COGNITO_ISSUER_TYPE=updated

Authentication Discovery for Agents

Publishes, at /.well-known/oauth-protected-resource, where clients obtain a token, and points every 401 Unauthorized at that document. An AI agent — or any MCP client — can then authenticate against this deployment without having been configured for it first. See Authentication Discovery for Agents for the full flow.

Nothing is published until OAUTH_RESOURCE_IDENTIFIER is set. With an Amazon Cognito user pool configured, that variable is the only one to set: the pool already names the issuer and the scopes, and both are published from it. The document is public and unauthenticated, since a client reads it before it has any credential.

OAUTH_RESOURCE_IDENTIFIER

Purpose : Public URL clients use to reach this deployment, published as the identity of the protected resource

Default : None — no discovery document is published, and 401 responses only state that a bearer token is expected

Requirement : Must be the exact origin clients dial — scheme and host, an explicit port only when it is not the default one for the scheme, and no path, query or fragment. Clients compare it character by character against the URL they used, so https://api.example.com and https://api.example.com:443 are not interchangeable. Requires OAUTH_AUTHORIZATION_SERVERS, unless a user pool supplies the issuer.

export OAUTH_RESOURCE_IDENTIFIER=https://api.example.com

OAUTH_AUTHORIZATION_SERVERS

Purpose : Issuer URLs of the OAuth 2.0 authorization servers that issue tokens for this deployment, comma-separated

Default : The issuer of the AWS_COGNITO_USER_POOL_ID pool, when one is configured — otherwise none

Effect : A client reads each issuer's own metadata to find where to sign in, so this deployment never describes the sign-in flow itself. A load balancer or API gateway authenticating in front of stdapi.ai publishes the issuer of whichever provider it uses.

With a user pool configured, leave this unset: the pool issues the tokens the deployment accepts, so its own issuer is published — `https://cognito-idp.<region>.amazonaws.com/<pool-id>`, or `https://issuer-cognito-idp.<region>.amazonaws.com/<pool-id>` when [`AWS_COGNITO_ISSUER_TYPE`](#aws-cognito-issuer-type) is `updated`. `<region>` and `<pool-id>` come from the pool ID itself, and the host follows the pool Region's AWS partition (`amazonaws.com.cn` in China, `amazonaws.eu` in the European Sovereign Cloud). Set the variable only to publish further issuers.

Requirement : Each entry is an https URL with no query or fragment. Required when OAUTH_RESOURCE_IDENTIFIER is set and no user pool is configured. When one is, the list must include the pool's own issuer — a client sent anywhere else obtains a token every request refuses, so startup fails instead.

export OAUTH_AUTHORIZATION_SERVERS=https://cognito-idp.eu-west-3.amazonaws.com/eu-west-3_a1b2c3d4e

OAUTH_SCOPES_SUPPORTED

Purpose : Scopes a token needs to call this API, comma-separated

Default : AWS_COGNITO_REQUIRED_SCOPES — with neither set, no scope is advertised and a client asks for whatever its own configuration names

Effect : Advertised both in the discovery document and in the 401 challenge, so a client asks its authorization server for the right scopes on its first attempt. The scopes a token must carry to be accepted are exactly the scopes to ask for, so they are published unless this variable names others.

export OAUTH_SCOPES_SUPPORTED=stdapi/invoke

API Compatibility

Configure the base URL paths for OpenAI and Anthropic-compatible API routes.

OPENAI_ROUTES_PREFIX

Purpose : Base path prefix for OpenAI-compatible API routes

Default : `` (empty, routes mounted at root)

Requirement : Empty, or a path starting with / with no trailing slash, using only alphanumeric characters and . _ ~ - per segment; must differ from ANTHROPIC_ROUTES_PREFIX and COHERE_ROUTES_PREFIX

Effect : All OpenAI-compatible endpoints will be mounted under this prefix

export OPENAI_ROUTES_PREFIX=/api

Example Endpoints

With the prefix /api, endpoints are available at:

  • /api/v1/chat/completions
  • /api/v1/models
  • /api/v1/embeddings

ANTHROPIC_ROUTES_PREFIX

Purpose : Base path prefix for Anthropic-compatible API routes

Default : /anthropic

Requirement : A path starting with / with no trailing slash, using only alphanumeric characters and . _ ~ - per segment; must differ from OPENAI_ROUTES_PREFIX and COHERE_ROUTES_PREFIX

Effect : All Anthropic-compatible endpoints will be mounted under this prefix

export ANTHROPIC_ROUTES_PREFIX=/anthropic

Example Endpoints

With the default prefix /anthropic, endpoints are available at:

  • /anthropic/v1/messages

Custom Prefix

You can change the prefix to match your organization's API structure:

export ANTHROPIC_ROUTES_PREFIX=/api/anthropic

This would mount the Messages API at /api/anthropic/v1/messages

COHERE_ROUTES_PREFIX

Purpose : Base path prefix for Cohere-compatible API routes

Default : /cohere

Requirement : A path starting with / with no trailing slash, using only alphanumeric characters and . _ ~ - per segment; must differ from OPENAI_ROUTES_PREFIX and ANTHROPIC_ROUTES_PREFIX

Effect : All Cohere-compatible endpoints will be mounted under this prefix

export COHERE_ROUTES_PREFIX=/cohere

Example Endpoints

With the default prefix /cohere, endpoints are available at:

  • /cohere/v2/rerank

CORS Configuration

Configure Cross-Origin Resource Sharing (CORS) to control which web origins can access your API from browsers.

CORS_ALLOW_ORIGINS

Purpose : List of origins allowed to make cross-origin requests

Format : JSON array of origin URLs

Default : None (CORS not enabled)

Best Practice : Only enable if your API is accessed from web browsers; specify exact origins in production

# Not configured (default) - CORS middleware not enabled
# Browser cross-origin requests will be blocked
# No environment variable needed

# Development: Allow all origins
export CORS_ALLOW_ORIGINS='["*"]'

# Production: Specific origins only
export CORS_ALLOW_ORIGINS='["https://myapp.com", "https://app.example.com"]'

# Multiple environments
export CORS_ALLOW_ORIGINS='["https://app.example.com", "https://staging.example.com"]'

What is CORS?

Cross-Origin Resource Sharing (CORS) is a browser security mechanism that restricts web pages from making requests to a different domain than the one serving the web page.

Without CORS enabled:

  • Browser requests from web applications will fail due to missing CORS headers
  • Non-browser clients (curl, SDKs, mobile apps, server-to-server) work normally
  • Most secure default - no cross-origin access from browsers

With CORS enabled:

  • Browsers can make requests from allowed origins
  • Preflight OPTIONS requests are handled automatically
  • Non-browser clients continue to work normally

Security Consideration

  • Default (not configured): CORS is disabled. Browser cross-origin requests will fail. This is the most secure default.
  • ["*"]: Allows requests from any web origin. Convenient for development but not recommended for production.
  • Specific origins: Only allows requests from listed origins. Recommended for production.

CORS Behavior

  • When CORS_ALLOW_ORIGINS is not configured (default), CORS is not enabled
  • When configured with specific origins or ["*"], CORS is enabled with:
    • Authorization headers with credentials allowed
    • All HTTP methods allowed
    • All request headers allowed

When to Configure

Configure CORS_ALLOW_ORIGINS when:

  • Your API is accessed from browser-based web applications (React, Vue, Angular, etc.)
  • Building a web frontend that calls your API from a different domain
  • Developing locally with web apps (browser at localhost:3000 calling API at localhost:8000)

When NOT to Configure

Do not configure CORS when:

  • Your API is only accessed from server-to-server integrations
  • Your API is only accessed from mobile apps or desktop clients
  • Your API is only accessed from CLI tools or SDKs
  • Your API is only accessed from non-browser HTTP clients

Non-browser clients don't enforce CORS, so enabling it is unnecessary overhead.


Trusted Host Configuration

Configure Host header validation to protect against Host header injection attacks.

TRUSTED_HOSTS

Purpose : List of trusted Host header values for validation

Format : JSON array of hostnames (supports wildcards)

Default : None (no Host header validation)

Best Practice : Use AWS ALB host-based routing rules instead when possible for better performance and management

# Production: Specific hosts only
export TRUSTED_HOSTS='["api.example.com", "www.example.com"]'

What is Host Header Validation?

The Host header in HTTP requests specifies the domain name of the server. Validating it prevents Host header injection attacks (manipulated Host headers used to poison caches or exploit application logic) and web cache poisoning.

Security Consideration: prefer ALB host-based routing

Configure AWS ALB listener rules to validate the Host header and forward traffic only for approved hostnames — this rejects bad requests at the load balancer, before they reach the application, and is centrally managed. See the example below.

Use TRUSTED_HOSTS only when you can't configure host-based routing at the load balancer level (no ALB, or you need application-level defense-in-depth).

Wildcard Support

  • *.example.com matches any subdomain (api.example.com, app.example.com, ...)
  • example.com matches only the exact domain
  • * matches all hosts — not recommended, equivalent to no validation

Common Configurations

Multi-Domain with Subdomains:

export TRUSTED_HOSTS='["*.example.com", "*.myapp.com", "api.production.com"]'

Development and Production:

export TRUSTED_HOSTS='["api.example.com", "localhost", "127.0.0.1"]'

Host Validation Behavior

  • Not configured (default): Host header validation is not enabled
  • Configured: requests with a non-matching Host header are rejected with HTTP 400 Bad Request

Container health probe

Validation applies to /health like any other path, so the container image's HEALTHCHECK derives its Host header from this setting: it requests /health on 127.0.0.1:$GRANIAN_PORT announcing the first entry of TRUSTED_HOSTS. * or an unset value becomes localhost, and a leading *. becomes healthcheck. (so *.example.com is probed as healthcheck.example.com).

A correct list therefore keeps the container healthy with no extra entry to add. Do not replace the probe with a hand-written curl call in a Compose healthcheck: block or an ECS task definition healthCheck: it would send an untrusted Host and get a 400.

Load balancer health checks are rejected by default

An ALB or NLB target-group health check does not send your domain name: it addresses the target directly, so the Host header carries the target's IP address. With TRUSTED_HOSTS set to domain names, every one of those probes gets HTTP 400, the target never turns healthy, and the load balancer serves 503 — a failure that looks like a broken deployment rather than a configuration choice.

Target-group health-check settings offer no Host header override, so either keep the Host allow-list at the load balancer (the recommended option above, leaving TRUSTED_HOSTS unset) or make sure the address the health check actually sends is in the list.

AWS ALB Host-Based Routing Example

Via AWS Console: EC2 → Load Balancers → Your ALB → Listeners → add a rule on the HTTPS (443) listener with condition "Host header" is api.example.com, forwarding to the target group only on match.

Via AWS CLI:

aws elbv2 create-rule \
  --listener-arn arn:aws:elasticloadbalancing:... \
  --priority 1 \
  --conditions Field=host-header,Values=api.example.com \
  --actions Type=forward,TargetGroupArn=arn:aws:elasticloadbalancing:...

Benefits: rejected at the load balancer (better performance, reduced load on application servers), centralized policy management, and ALB metrics/logging for rejected requests.


Proxy Headers Configuration

Configure X-Forwarded-* header processing when running behind reverse proxies or load balancers.

ENABLE_PROXY_HEADERS

Purpose : Enable trusting X-Forwarded-* headers from reverse proxies

Type : Boolean

Default : false (disabled)

Best Practice : Only enable when running behind a trusted reverse proxy

# Disabled (default) - do not trust X-Forwarded-* headers
# No environment variable needed

# Enable when behind reverse proxy
export ENABLE_PROXY_HEADERS=true

What are X-Forwarded Headers?

When your application runs behind a reverse proxy (nginx, Apache, AWS ALB, CloudFront, etc.), the proxy sits between clients and your application. Without proxy header processing:

  • The application sees the proxy's IP address instead of the client's real IP
  • The application sees the proxy-to-app connection (e.g., HTTP) instead of the original client connection (e.g., HTTPS)
  • The application cannot distinguish between different clients behind the proxy

Reverse proxies add X-Forwarded-* headers to preserve the original request information:

  • X-Forwarded-For - Client's real IP address (and chain of proxies)
  • X-Forwarded-Proto - Original protocol (http/https)
  • X-Forwarded-Port - Original port number

Security Warning

CRITICAL: Only enable ENABLE_PROXY_HEADERS when running behind a trusted reverse proxy that properly sets X-Forwarded-* headers.

If enabled without a trusted proxy:

  • Clients can spoof their IP address by sending fake X-Forwarded-For headers
  • Security controls based on client IP (rate limiting, allowlists) can be bypassed
  • Logging and monitoring will record incorrect client information
  • Authentication and authorization decisions may be affected

Never enable this setting if your application is directly exposed to the internet without a reverse proxy.

Common Deployment Scenarios

Scenario 1: Direct to Internet (No Proxy)

# Do NOT enable proxy headers
# ENABLE_PROXY_HEADERS should remain false (default)

Your application receives requests directly from clients.

Scenario 2: Behind AWS ALB/CloudFront

export ENABLE_PROXY_HEADERS=true

AWS load balancer or CDN forwards requests to your application.

Scenario 3: Multiple AWS Proxy Layers

export ENABLE_PROXY_HEADERS=true

Example: CloudFront → ALB → Your Application

Proxy Headers Behavior

  • When ENABLE_PROXY_HEADERS is false (default), X-Forwarded- headers are not trusted*
  • When enabled, the server processes X-Forwarded-For, X-Forwarded-Proto, and X-Forwarded-Port headers to determine client information
  • Which peers' headers are trusted is controlled by PROXY_TRUSTED_HOSTS — the default * trusts every peer, so restrict it to your reverse proxy's IP range

When to Enable

Enable ENABLE_PROXY_HEADERS when:

  • Deployed behind AWS ALB, NLB, API Gateway, or CloudFront
  • Running behind any reverse proxy that sets X-Forwarded-* headers

AWS Proxy Configuration

AWS ALB, NLB, and CloudFront automatically set X-Forwarded-* headers - no additional configuration needed.

When you enable ENABLE_PROXY_HEADERS=true, your application will trust these headers to determine:

  • Client's real IP address (from X-Forwarded-For)
  • Original protocol (from X-Forwarded-Proto: http/https)
  • Original port (from X-Forwarded-Port)

PROXY_TRUSTED_HOSTS

Purpose : Restrict which peer IPs may set trusted X-Forwarded-* headers when ENABLE_PROXY_HEADERS is enabled

Type : JSON array of IPs/CIDRs, or *

Default : * (trust every peer — backward compatible)

Best Practice : Restrict to your reverse proxy's IP range so direct clients cannot spoof X-Forwarded-For

# Trust forwarded headers only from the VPC / proxy range
export ENABLE_PROXY_HEADERS=true
export PROXY_TRUSTED_HOSTS='["10.0.0.0/8"]'

Only effective with ENABLE_PROXY_HEADERS=true

This setting has no effect unless ENABLE_PROXY_HEADERS is enabled. With the default *, any client that can reach the server directly can forge X-Forwarded-For, poisoning the client IP recorded in logs and OpenTelemetry spans. Restrict it to the address range of your load balancer or reverse proxy (AWS ALB/CloudFront, nginx, etc.).

Configured automatically by the official Terraform module

The stdapi-ai Terraform module sets this for you when the ALB is enabled with client IP logging (alb_enabled = true, log_client_ip = true): it enables proxy headers and pins PROXY_TRUSTED_HOSTS to the ALB's subnet CIDRs, so only the load balancer is trusted and direct clients cannot forge X-Forwarded-For. Override it with the module's proxy_trusted_hosts variable when fronting the ALB with an additional proxy (for example CloudFront).

On a dual-stack listener, cover the IPv4-mapped form too

With GRANIAN_HOST=:: the operating system reports an IPv4 peer as an IPv4-mapped IPv6 address such as ::ffff:10.0.1.5, which belongs to no IPv4 network and therefore matches no IPv4 entry here. Add the mapped range alongside the plain one — an IPv4 /16 becomes a /112 once the 96-bit mapping prefix is counted:

export PROXY_TRUSTED_HOSTS='["10.0.0.0/16", "::ffff:10.0.0.0/112"]'

Miss it and the proxy stops being trusted: X-Forwarded-For is ignored and the load balancer's own address is recorded as the client IP. The Terraform module derives these entries for you, including for values passed to its proxy_trusted_hosts variable.


TLS / SSL Configuration

Configure end-to-end TLS encryption within the container. These are native Granian environment variables and are available with the provided container images.

GRANIAN_SSL_CERTIFICATE

Purpose : Path to the SSL certificate file

Type : File path

GRANIAN_SSL_KEYFILE

Purpose : Path to the SSL private key file (PKCS#8 format only)

Type : File path

GRANIAN_SSL_KEYFILE_PASSWORD

Purpose : Password for the private key file

Type : String

GRANIAN_SSL_PROTOCOL_MIN

Purpose : Minimum supported TLS version (tls1.2 or tls1.3)

Type : Enum

Default : tls1.3

GRANIAN_SSL_CA

Purpose : Path to the CA certificate bundle used to verify client certificates (mTLS)

Type : File path

GRANIAN_SSL_CLIENT_VERIFY

Purpose : Enable client certificate verification (mTLS)

Type : Boolean

Default : false (disabled)


GZip Compression

Configure automatic GZip compression for HTTP responses to reduce bandwidth usage and improve response times.

ENABLE_GZIP

Purpose : Enable GZip compression for HTTP responses

Type : Boolean

Default : false (disabled)

Best Practice : Use AWS ALB or CloudFront compression instead when available for better performance

# Disabled (default) - no response compression
# No environment variable needed

# Enable GZip compression (responses larger than 1 KiB will be compressed)
export ENABLE_GZIP=true

How GZip Compression Works

When enabled, the server automatically:

  1. Checks if the response size exceeds 1 KiB (1024 bytes)
  2. Verifies the client supports compression (via Accept-Encoding: gzip header)
  3. Compresses the response body using gzip
  4. Adds Content-Encoding: gzip header to the response

Typical compression ratios for JSON responses: 60-80% size reduction

Recommended: Use AWS Compression Services

Instead of enabling application-level compression, enable compression at the AWS layer — it offloads the CPU cost from your application servers, at the price of managing it in AWS instead of a single environment variable:

  • AWS ALB — enable the compression.enabled target group attribute (documentation)
  • Amazon CloudFront — enable "Compress Objects Automatically" in the distribution behavior settings (documentation)

When to Enable Application-Level Compression

Enable ENABLE_GZIP only when:

  • You're not using AWS ALB or CloudFront
  • Your API returns large JSON responses and you want to reduce bandwidth
  • Local development or non-AWS deployments

When NOT to Enable

Do not enable when:

  • You're behind AWS ALB with compression enabled
  • You're using CloudFront with compression enabled
  • CPU usage is a concern (compression adds CPU overhead)

Enabling compression at multiple layers is redundant and wastes CPU resources.

Compression Behavior

  • When ENABLE_GZIP is false (default), compression is not enabled
  • When enabled, only responses meeting these criteria are compressed:
    • Response size ≥ 1 KiB (1024 bytes)
    • Client sends Accept-Encoding: gzip header
    • Response does not already have Content-Encoding header
  • Streaming responses are compressed on-the-fly

MCP (Model Context Protocol)

When enabled, stdapi.ai exposes its API endpoints as MCP tools, allowing AI clients and agents to call them directly using the Model Context Protocol. The full list of available tool names is documented in API Overview → MCP Tools.

Both transport types can be enabled independently or simultaneously.

ENABLE_MCP_STREAMABLE_HTTP

Purpose : Enable the MCP server using Streamable HTTP transport — the recommended method

Type : Boolean

Default : false

Behavior : Exposes an MCP-compatible endpoint at /mcp. AI clients connect using standard HTTP requests following the MCP Streamable HTTP specification.

# Disabled (default)
# No environment variable needed

# Enable MCP Streamable HTTP transport
export ENABLE_MCP_STREAMABLE_HTTP=true

MCP_STATELESS_HTTP

Purpose : Serve the Streamable HTTP transport without server-side sessions

Type : Boolean

Default : false

Behavior : Each request to /mcp is handled by a fresh transport that keeps no state. Clients may call tools/list and tools/call without an initialize handshake, an Mcp-Session-Id the server never issued is accepted rather than rejected, and any replica may serve any request.

Requires : ENABLE_MCP_STREAMABLE_HTTP=true. Ignored otherwise.

# Sessions enabled (default)
# No environment variable needed

# Stateless transport
export ENABLE_MCP_STREAMABLE_HTTP=true
export MCP_STATELESS_HTTP=true

ENABLE_MCP_SSE

Purpose : Enable the MCP server using Server-Sent Events (SSE) transport

Type : Boolean

Default : false

Behavior : Exposes MCP endpoints at /sse for AI clients that require the SSE transport protocol.

# Disabled (default)
# No environment variable needed

# Enable MCP SSE transport
export ENABLE_MCP_SSE=true

Transport Recommendation

HTTP transport (ENABLE_MCP_STREAMABLE_HTTP) is the recommended method. It implements the latest MCP Streamable HTTP specification and provides better session management and more robust connection handling.

SSE transport (ENABLE_MCP_SSE) is maintained for backwards compatibility with older MCP client implementations. Prefer HTTP for new deployments.

Both transports can be enabled simultaneously to support clients with different requirements:

export ENABLE_MCP_STREAMABLE_HTTP=true
export ENABLE_MCP_SSE=true

The MCP server card (/.well-known/mcp/server-card.json) declares a single transport: Streamable HTTP (/mcp) whenever it is enabled, otherwise SSE (/sse). When both are enabled, /sse is therefore not listed in the card, but it remains fully functional for clients configured with it explicitly.

MCP_INCLUDE_TOOLS

Purpose : Expose only a specific subset of MCP tools; all others are hidden

Format : Comma-separated list of tool names (duplicates are automatically removed)

Default : None (all tools exposed)

# All tools exposed by default
# No environment variable needed

# Expose only specific tools
export MCP_INCLUDE_TOOLS="openai_chat_completion,openai_embedding,search_models"

# When both MCP_INCLUDE_TOOLS and MCP_EXCLUDE_TOOLS are specified,
# tools in MCP_EXCLUDE_TOOLS are removed from MCP_INCLUDE_TOOLS:
export MCP_INCLUDE_TOOLS="openai_chat_completion,openai_embedding,search_models"
export MCP_EXCLUDE_TOOLS="openai_files_delete,anthropic_files_delete"
# Result: only openai_chat_completion, openai_embedding, search_models are exposed

See API Overview → MCP Tools for the full list of available tool names.

Token Usage for Complex API Tools

anthropic_message, openai_chat_completion, and openai_response map to large, complex APIs that may use many tokens (prompt, completion, and tool definitions). Select these tools only if your workflow requires the full API capabilities.

MCP_EXCLUDE_TOOLS

Purpose : Hide specific MCP tools from clients; all others remain exposed

Format : Comma-separated list of tool names (duplicates are automatically removed)

Default : None (no tools excluded)

Behavior with MCP_INCLUDE_TOOLS

When both MCP_INCLUDE_TOOLS and MCP_EXCLUDE_TOOLS are specified, tools in MCP_EXCLUDE_TOOLS are removed from MCP_INCLUDE_TOOLS. The remaining tools in MCP_INCLUDE_TOOLS are what get exposed.

# No tools excluded by default
# No environment variable needed

# Exclude destructive tools
export MCP_EXCLUDE_TOOLS="openai_files_delete,anthropic_files_delete"

See API Overview → MCP Tools for the full list of available tool names.

Tool Selection Best Practices

stdapi.ai exposes a fixed set of tools derived from its API surface — you can include or exclude them by name, but cannot modify or rename them. See API Overview → MCP Tools for the full catalog.

Start from the minimum, not the maximum

By default all tools are exposed. It is safer and more effective to begin with a narrow MCP_INCLUDE_TOOLS list covering only what the workflow needs, then expand it deliberately. LLMs perform better with fewer choices, and many AI providers cap the number of active tools per session.

Always include search_models for agent model discovery

search_models is the recommended tool for agents to discover available model IDs — it supports capability-based filtering (by modality, route, region, streaming support) and returns richer metadata than openai_model_list or anthropic_model_list. Include it in every agent configuration so the agent can resolve the right model dynamically rather than relying on hardcoded IDs:

export MCP_INCLUDE_TOOLS="openai_chat_completion,search_models,openai_embedding"

Always exclude file deletion tools unless required

Uploaded files are the only durable, stateful data managed by stdapi.ai — deletion is permanent and cannot be undone. Unless your workflow explicitly needs to delete files, always suppress these tools:

export MCP_EXCLUDE_TOOLS="openai_files_delete,anthropic_files_delete"

Exclude high-cost tools unless the workflow requires them

Image generation (openai_image_generation, openai_image_edit, openai_image_variation) and speech synthesis (openai_audio_speech) incur a per-call cost that accumulates quickly if an agent invokes them speculatively. Only include them when the use case calls for it and the agent's decision to generate images or audio is intentional.

Use MCP_INCLUDE_TOOLS for the tightest control

For predictable, well-defined workflows, listing tools explicitly with MCP_INCLUDE_TOOLS is more reliable than maintaining an exclusion list. For example, a workflow limited to text generation and model discovery needs only:

export MCP_INCLUDE_TOOLS="openai_chat_completion,search_models"

Note

Health and metadata endpoints are never exposed as MCP tools, so they do not need to be listed in MCP_EXCLUDE_TOOLS.


SSRF Protection

Configure Server-Side Request Forgery (SSRF) protection to prevent unauthorized access to internal networks.

SSRF_PROTECTION_BLOCK_PRIVATE_NETWORKS

Purpose : Enable SSRF protection by blocking requests to private/local networks

Type : Boolean

Default : true (enabled for security)

Best Practice : Keep enabled in production to protect against SSRF attacks

# Enabled (default) - block private networks
# No environment variable needed

# Disable only in controlled environments that need local network access
export SSRF_PROTECTION_BLOCK_PRIVATE_NETWORKS=false

What is SSRF Protection?

Server-Side Request Forgery (SSRF) is an attack where an attacker can make the server send requests to unintended destinations, including internal network resources.

SSRF protection has two layers:

  1. Baseline Protection (Always Enabled) - Cannot be disabled:

    • :material-loopback: Loopback Addresses - 127.0.0.0/8, ::1
    • Unspecified Addresses - 0.0.0.0, ::
    • Link-Local Addresses - 169.254.0.0/16, fe80::/10
    • Reserved IP Ranges - IETF reserved addresses
    • Multicast Addresses - Multicast IP ranges
  2. Private Network Protection (Controlled by this setting):

    • Every Non-Globally-Reachable Address - anything outside the public Internet address space, in both families and in IPv4-mapped IPv6 form
    • Examples - RFC 1918 (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16), IPv6 unique local (fc00::/7), RFC 6598 shared address space (100.64.0.0/10), benchmarking (198.18.0.0/15) and documentation ranges

Security Warning

CRITICAL: Only disable SSRF_PROTECTION_BLOCK_PRIVATE_NETWORKS in controlled environments where accessing internal networks is explicitly required and safe.

If disabled, private network protection is removed:

  • Attackers may be able to reach any non-globally-reachable address (private networks, shared address space, and the other special-purpose ranges) through your API
  • Internal services on private networks (databases, admin panels, internal APIs) may be exposed
  • Internal APIs without authentication may be exploited

Important: Even when disabled, baseline protection remains active and prevents access to:

  • Loopback addresses (127.0.0.1, localhost) - always blocked
  • Link-local addresses (169.254.x.x) including AWS EC2 metadata endpoint - always blocked
  • Reserved and multicast addresses - always blocked

When to Disable

Disable SSRF_PROTECTION_BLOCK_PRIVATE_NETWORKS only when:

  • Your application legitimately needs to access internal network resources
  • Local development environment where accessing localhost services is required
  • You have other security controls in place (network segmentation, firewall rules)
  • Running in isolated Docker/container environments with restricted network access

Defense in Depth

Even with SSRF protection enabled, implement additional security measures:

  • Network Segmentation - Isolate application servers from sensitive internal networks
  • :material-firewall: Firewall Rules - Restrict outbound connections from application servers
  • Security Groups - Use AWS security groups to limit network access
  • Monitoring - Log and monitor outbound requests for suspicious patterns

Request Limits

Bound per-request resource usage to protect the server when the API is exposed to untrusted clients.

MAX_INPUT_FILE_SIZE

Purpose : Cap the size of an inline input file loaded into memory to protect against memory-exhaustion (DoS)

Type : Integer (bytes)

Default : 0 (disabled — no limit)

Best Practice : Set a limit aligned with your largest expected inline input (e.g. 26214400 for 25 MiB) when the API is exposed to untrusted clients

# Disabled (default) - no size limit
# No environment variable needed

# Reject inline inputs larger than 25 MiB
export MAX_INPUT_FILE_SIZE=26214400

What is limited

The limit applies to file content that is loaded into memory for model input:

Requests exceeding the limit are rejected with HTTP 413 before the content is fully decoded or downloaded. For downloads, the body is streamed and aborted as soon as the limit is exceeded, so a spoofed Content-Length cannot bypass it.

Streaming uploads are not affected, so large file transfers remain possible:

  • Multipart form uploads
  • Files API ingest from HTTP(S) URLs and S3-to-S3 copies

MAX_CONCURRENT_INPUT_DOWNLOADS

Purpose : Bound the number of input files fetched or resolved concurrently within a single request

Type : Integer (> 0)

Default : 8

Best Practice : Keep a modest value so a single request with many remote inputs cannot exhaust sockets/memory or amplify outbound requests against a target

# Allow up to 4 concurrent input downloads per request
export MAX_CONCURRENT_INPUT_DOWNLOADS=4

Behaviour

Each remote input (image, document, or audio referenced by URL or S3 URI) is fetched in parallel, capped at this many at a time. Excess inputs queue and run as slots free up, so requests still complete — they are only paced. This prevents a request carrying thousands of URLs from opening thousands of simultaneous connections (socket/memory exhaustion and SSRF amplification).


Observability (OpenTelemetry)

Configure distributed tracing for debugging and performance monitoring. stdapi.ai integrates with AWS X-Ray, Jaeger, DataDog, and other OTLP-compatible systems.

OTEL_ENABLED

Purpose : Enable or disable OpenTelemetry tracing

Type : Boolean

Default : false

export OTEL_ENABLED=true

Performance Consideration

Disable in performance-critical deployments where observability is not needed.

OTEL_SERVICE_NAME

Purpose : Service identifier in trace visualizations

Default : stdapi.ai

Best Practice : Use descriptive names with environment information

export OTEL_SERVICE_NAME=stdapi-production-us-east-1

OTEL_EXPORTER_ENDPOINT

Purpose : OTLP HTTP endpoint URL for sending traces

Default : http://127.0.0.1:4318/v1/traces

Protocol : Must support OTLP HTTP format

AWS X-Ray (via ADOT):

export OTEL_EXPORTER_ENDPOINT=http://127.0.0.1:4318/v1/traces

Jaeger:

export OTEL_EXPORTER_ENDPOINT=http://jaeger:14268/api/traces

Cloud Provider OTLP:

# Use provider-specific OTLP endpoints
export OTEL_EXPORTER_ENDPOINT=https://your-provider-otlp-endpoint.com/v1/traces

OTEL_SAMPLE_RATE

Purpose : Percentage of requests to trace (controls cost vs. observability)

Type : Float (0.0 to 1.0)

Default : 1.0 (100%)

Development:

# Trace everything for debugging
export OTEL_SAMPLE_RATE=1.0

Production (Moderate Traffic):

# Sample 10% of requests
export OTEL_SAMPLE_RATE=0.1

Production (High Traffic):

# Sample 1% of requests
export OTEL_SAMPLE_RATE=0.01

Sampling Recommendations

Sample Rate Use Case
1.0 (100%) Development, debugging, low-traffic services
0.1 (10%) Production with moderate traffic
0.01 (1%) High-traffic production services
0.0 (0%) Equivalent to disabling tracing

API Documentation Routes

stdapi.ai provides automatic API documentation routes, which are disabled by default for security in production environments.

Security Consideration

Exposing API documentation routes in production can reveal internal API structure, available endpoints, and request/response schemas to potential attackers. Only enable these routes in development/testing environments or when absolutely necessary.

Agent Discovery

The machine-readable API catalog at /.well-known/api-catalog (RFC 9727 Linkset) is always served, regardless of the settings below. Enabling a route adds its entry to the catalog:

  • ENABLE_OPENAPI_JSON — adds the service-desc link to /openapi.json
  • ENABLE_DOCS or ENABLE_REDOC — adds the service-doc link to /docs or /redoc (Swagger UI takes precedence when both are enabled)
  • ENABLE_MCP_STREAMABLE_HTTP or ENABLE_MCP_SSE — adds the mcp-server-card link

The same links are also advertised as RFC 8288 Link headers on the root endpoint (/). That header is only emitted when at least one of these routes is enabled; with all of them disabled, the catalog is still reachable but carries no links.

ENABLE_DOCS

Purpose : Enable interactive Swagger UI documentation at /docs

Type : Boolean

Default : false (disabled)

# Enable for development
export ENABLE_DOCS=true

Interactive Documentation Features

The /docs endpoint provides an interactive interface to:

  • Browse all available API endpoints
  • Test API requests directly from the browser
  • View request/response schemas
  • Understand parameter requirements

ENABLE_REDOC

Purpose : Enable ReDoc documentation UI at /redoc

Type : Boolean

Default : false (disabled)

# Enable for development
export ENABLE_REDOC=true

ReDoc Features

The /redoc endpoint provides a clean, responsive documentation interface with:

  • Three-panel layout for easy navigation
  • Enhanced schema visualization
  • Better rendering for complex APIs
  • Export to OpenAPI specification

Works with no outbound access

Both pages are served entirely by the gateway: the container image ships Swagger UI and ReDoc itself, pinned to an exact release and verified against a recorded SHA-256 during the build, alongside the icon and the schema. A browser that can reach the gateway renders them, with no request to any CDN, font host or other third party — so they work unchanged in an air-gapped VPC, behind an egress allow-list, or under a strict content security policy.

Static Documentation Available

ReDoc API documentation is also available as static documentation at API Reference without requiring this endpoint to be enabled.

ENABLE_OPENAPI_JSON

Purpose : Enable OpenAPI schema JSON endpoint at /openapi.json

Type : Boolean

Default : false (disabled)

# Enable for development
export ENABLE_OPENAPI_JSON=true

OpenAPI Schema

The /openapi.json endpoint provides the raw OpenAPI 3.0 specification, useful for:

  • Generating API clients in various languages
  • Import into API testing tools (Postman, Insomnia)
  • API documentation generation
  • Contract testing and validation

Automatic Enablement

If either ENABLE_DOCS or ENABLE_REDOC is set to true, the /openapi.json endpoint will be automatically enabled since both documentation UIs require the OpenAPI schema to function. You only need to explicitly set ENABLE_OPENAPI_JSON=true if you want to expose the schema endpoint without enabling the documentation UIs.

Development Configuration

Enable all documentation routes for local development:

export ENABLE_DOCS=true
export ENABLE_REDOC=true
# ENABLE_OPENAPI_JSON is automatically enabled when ENABLE_DOCS or ENABLE_REDOC is true

Or enable only Swagger UI:

export ENABLE_DOCS=true
# ENABLE_OPENAPI_JSON is automatically enabled

Or enable only ReDoc:

export ENABLE_REDOC=true
# ENABLE_OPENAPI_JSON is automatically enabled

Production Best Practice

# Keep all routes disabled in production (default)
# No environment variables needed - defaults to false

Production Warning

Never enable these routes in production unless you have specific security controls in place (e.g., IP allowlisting, VPN-only access, or additional authentication layer).


Validation and Logging

For comprehensive logging and monitoring information, see the Logging and Monitoring guide.

STRICT_INPUT_VALIDATION

Purpose : Reject API requests containing unknown/extra fields instead of ignoring them

Type : Boolean

Default : false

# Returns HTTP 400 for requests with unexpected fields
export STRICT_INPUT_VALIDATION=true

CHAT_COMPLETIONS_REASONING_FIELD

Purpose : Choose which field carries a reasoning model's thinking text on /v1/chat/completions

Type : String

Default : reasoning_content

Options : reasoning_content, reasoning, none

Behavior : The OpenAI Chat Completions API returns no thinking text of its own — it reports only a reasoning_tokens count — so the providers that do return it have settled on two different names. reasoning_content is the DeepSeek spelling, which most clients that read reasoning at all look for first. reasoning is the name used by OpenRouter and vLLM. none emits neither, keeping responses strictly OpenAI-shaped. : The setting applies to both the completed message and the streamed deltas, so a client never sees one name while streaming and another at the end. Callers can also suppress reasoning per request with include_reasoning: false or reasoning: {"exclude": true}, whatever this is set to.

# Default: the name most clients read
export CHAT_COMPLETIONS_REASONING_FIELD=reasoning_content

# For clients written against OpenRouter or vLLM
export CHAT_COMPLETIONS_REASONING_FIELD=reasoning

# Strict OpenAI shape: never return thinking text
export CHAT_COMPLETIONS_REASONING_FIELD=none

LOG_LEVEL

Purpose : Control the minimum severity of log events written to STDOUT

Default : info

Options : info, warning, error, critical, disabled

Behavior : Only log events at or above the configured level are output. Log levels are ordered by severity: info < warning < error < critical

# Default: Output all log events
export LOG_LEVEL=info

# Production: Suppress info logs, show only warnings and higher
export LOG_LEVEL=warning

# Critical only: Show only critical errors
export LOG_LEVEL=critical

# Disable logging: Suppress all log output (not recommended)
export LOG_LEVEL=disabled

Log Level Examples

Level Outputs Use Case
info info, warning, error, critical Development, debugging, full visibility
warning warning, error, critical Production (recommended for most deployments)
error error, critical High-traffic production, reduce log volume
critical critical only Minimal logging, only show fatal errors
disabled none Not recommended - disables all logging

Production Recommendation

For production deployments, warning is recommended to reduce log volume while maintaining visibility into issues. The info level can generate significant log volume in high-traffic environments.

For detailed information about log events, structure, and monitoring strategies, see the Logging and Monitoring guide.

LOG_REQUEST_PARAMS

Purpose : Include request and response parameters (JSON body, form, query) in logs for integration debugging

Type : Boolean

Default : false

# Enable for debugging (NOT recommended for production)
export LOG_REQUEST_PARAMS=true

Security and Cost Warning

Enabling LOG_REQUEST_PARAMS may expose sensitive data in logs. Use only in development/debugging environments.

Logging full request/response payloads can also significantly increase log ingestion and storage costs, especially for large LLM prompts, tool calls, and generated outputs. If you must enable it, prefer short log retention, targeted sampling, and temporary use only.

LOG_CLIENT_IP

Purpose : Enable logging of client IP addresses for each request and add IP to OpenTelemetry spans

Type : Boolean

Default : false (disabled for privacy)

# Disabled (default) - no client IP logging
# No environment variable needed

# Enable client IP logging
export LOG_CLIENT_IP=true

Client IP Behavior

When enabled, client IP addresses are:

  • Included in log output for each request
  • Added as the client.address attribute to OpenTelemetry spans (when OTEL_ENABLED=true)

The IP address depends on your proxy configuration:

With ENABLE_PROXY_HEADERS=true (behind reverse proxy):

  • Logs the real client IP address from the X-Forwarded-For header
  • Shows the actual end-user IP, not the proxy IP
  • Requires your reverse proxy (ALB, CloudFront, etc.) to set the header correctly

With ENABLE_PROXY_HEADERS=false (default):

  • Logs the direct connection IP address
  • Typically shows your reverse proxy or load balancer IP, not the end-user IP
  • Limited usefulness unless application is directly exposed to clients

When to Enable

Enable LOG_CLIENT_IP when:

  • You need client IP addresses for security auditing or compliance
  • Analyzing traffic patterns and geographic distribution
  • Investigating abuse, fraud, or suspicious activity
  • Debugging client-specific issues

Important: Also enable ENABLE_PROXY_HEADERS=true when behind AWS ALB, CloudFront, or other reverse proxies to log the real client IP instead of the proxy IP.

Privacy Consideration

Client IP addresses are considered personal data under privacy regulations like GDPR. When logging IP addresses:

  • Consider shorter log retention periods
  • Document the purpose in your privacy policy
  • Ensure logs are stored securely
  • Implement log deletion procedures aligned with your data retention policy

Configuration for AWS Deployments

Behind AWS ALB or CloudFront:

# Enable proxy headers to get real client IPs
export ENABLE_PROXY_HEADERS=true
# Enable client IP logging
export LOG_CLIENT_IP=true

Direct exposure (not recommended for production):

# Only enable client IP logging
export LOG_CLIENT_IP=true
# ENABLE_PROXY_HEADERS remains false (default)

TIMEZONE

Purpose : IANA timezone identifier used for request date and time

Type : String (IANA timezone identifier)

Default : UTC

# UTC (default)
export TIMEZONE=UTC

# North America
export TIMEZONE=America/New_York

# Europe
export TIMEZONE=Europe/London

CloudWatch Metrics and Cost Tracking

The behavior of these settings — EMF line structure, cost log format, pricing accuracy, regional price fallback, known limitations, and the price override format with examples — is documented in CloudWatch Metrics (EMF) and Cost Tracking in the Logging and Monitoring guide.

CLOUDWATCH_METRICS

Purpose : Emit per-request AWS-billed usage as CloudWatch Embedded Metric Format (EMF) log lines

Type : Boolean

Default : false

export CLOUDWATCH_METRICS=true

CLOUDWATCH_METRICS_NAMESPACE

Purpose : CloudWatch namespace under which the emitted usage metrics are grouped

Type : String

Default : stdapi

Requirement : 1-255 characters, alphanumeric plus . - _ / # :, must not start with the reserved AWS/ prefix

export CLOUDWATCH_METRICS_NAMESPACE=my-app-metrics

COST_TRACKING

Purpose : Enable real-time cost computation from live AWS pricing (details and accuracy caveats). Disabled by default: it requires the extra pricing:GetProducts IAM permission — see Cost Tracking IAM Permissions.

Type : Boolean

Default : false

export COST_TRACKING=true

COST_PRICE_OVERRIDES

Purpose : Operator-supplied unit price overrides for models not covered by the AWS Price List API (format and example)

Type : JSON object — keys are model IDs, values are dicts mapping dimension name to price per one unit

Default : {}


Bedrock Guardrails

Amazon Bedrock Guardrails add content filtering and safety controls to model inputs and outputs. The configured guardrail also powers the OpenAI-compatible Moderations API (POST /v1/moderations); without one, that API falls back to inline guardrail checks in supported regions, then Amazon Comprehend.

Configuration Options

Guardrails can be configured in three ways:

  1. Global - Via environment variables
  2. Per-request - Via HTTP headers
  3. Request body - Via amazon-bedrock-guardrailConfig object

Route Coverage

The configured guardrail applies to every route that serves a request directly. The Batch API is the exception: Amazon Bedrock batch inference cannot apply a guardrail, so a batch a configured guardrail would cover is refused rather than run unchecked. Chat routes use the native Bedrock integration; routes whose AWS backend has no guardrail mechanism enforce it through the ApplyGuardrail API: client-supplied text is checked as INPUT before the backend call and generated text as OUTPUT after it. On the Realtime API, where content streams both ways for as long as the session is open, the check runs per turn and an intervention ends the session — see Guardrail coverage for what that does and does not catch.

Routes Mechanism Checked content
Chat Completions, Responses, Completions, Anthropic Messages Native (Converse guardrailConfig / InvokeModel) Model input and output
Moderations ApplyGuardrail (classification) Submitted text and images
Embeddings (OpenAI and Cohere v1/v2) ApplyGuardrail INPUT — each text input
Rerank (Cohere v1/v2) ApplyGuardrail INPUT — query and each document text
Images Generations / Edits ApplyGuardrail INPUT — prompt (Variations has no text to check)
Videos ApplyGuardrail INPUT — prompt
Audio Speech ApplyGuardrail INPUT — text to synthesize
Audio Transcriptions (including streaming) ApplyGuardrail OUTPUT — transcript
Audio Translations ApplyGuardrail OUTPUT — translated text
Realtime ApplyGuardrail INPUT — each written item before it reaches the model, and each transcribed caller turn; OUTPUT — each completed answer

Cost Tracking

AWS bills the guardrail on every route it applies to, but only the ApplyGuardrail-enforced ones report the units consumed. The mechanism a route uses therefore decides whether its guardrail cost is visible.

Mechanism Guardrail cost in usage logs
ApplyGuardrail Tracked — the response returns the units each policy consumed
Native (Converse / InvokeModel) Not tracked — the response reports no guardrail units

On ApplyGuardrail routes, the units AWS reports appear as text_units and input_images under one amazon.bedrock-runtime-guardrail-* model per applied policy, each priced at that policy's own rate; see Moderations billing. A route that checks both INPUT and OUTPUT calls the API twice, so it records two sets of units for one request.

On native routes the guardrail still runs and AWS still charges for it, but the Converse and InvokeModel responses carry no unit counts for the gateway to record. Reported costs on these routes are lower than the AWS bill by the guardrail's share. Deriving the units from text length instead would be a guess, not a measurement, so none is made.

Intervention Behavior

On ApplyGuardrail-enforced routes, a blocking intervention fails the request with HTTP 400 and error code content_filter (the same code chat routes report as their finish reason), carrying the guardrail's configured blocked messaging. A masking-only intervention (sensitive-information anonymization) substitutes the masked text — input masking reaches the backend model, and a masked transcript or translation is returned on the plain json/text formats. Response formats that cannot carry masked text (srt, vtt, verbose_json, diarized_json) fail with the same content_filter error instead of leaking the unmasked content.

Global Configuration

AWS_BEDROCK_GUARDRAIL_IDENTIFIER

Purpose : ID of the Bedrock Guardrail to apply

Required : Yes (together with AWS_BEDROCK_GUARDRAIL_VERSION)

export AWS_BEDROCK_GUARDRAIL_IDENTIFIER=abc123def456

AWS_BEDROCK_GUARDRAIL_VERSION

Purpose : Version of the Bedrock Guardrail

Required : Yes (together with AWS_BEDROCK_GUARDRAIL_IDENTIFIER)

export AWS_BEDROCK_GUARDRAIL_VERSION=1

AWS_BEDROCK_GUARDRAIL_TRACE

Purpose : Trace level for guardrail evaluation

Options : disabled, enabled, enabled_full

Default : None (optional)

export AWS_BEDROCK_GUARDRAIL_TRACE=enabled

AWS_BEDROCK_ALLOW_GUARDRAIL_OVERRIDE

Purpose : Control whether users can override the global guardrail configuration at request level via HTTP headers

Default : false (disabled for security)

Security Consideration : When set to false (default) and a global guardrail is configured, only the global configuration is enforced, preventing users from bypassing or modifying safety controls. Set to true if you need to allow per-request guardrail customization to override the global configuration.

Auto-Enable Behavior : If no guardrail is configured at all — both AWS_BEDROCK_GUARDRAIL_IDENTIFIER and AWS_BEDROCK_GUARDRAIL_VERSION unset, and no model alias carrying one — this setting is automatically set to true at startup, allowing per-request guardrails when no policy is enforced.

export AWS_BEDROCK_ALLOW_GUARDRAIL_OVERRIDE=true

Per-Alias Guardrails

A model alias can carry its own guardrail, applied to the requests naming it and overriding the global one. That is how a single deployment publishes the same model under a strictly guarded name and an unguarded one.

Complete Guardrail Configuration

export AWS_BEDROCK_GUARDRAIL_IDENTIFIER=abc123def456
export AWS_BEDROCK_GUARDRAIL_VERSION=1
export AWS_BEDROCK_GUARDRAIL_TRACE=enabled
export AWS_BEDROCK_ALLOW_GUARDRAIL_OVERRIDE=false  # Default: prevent overrides

Per-Request Guardrail Configuration

Header Usage Behavior

Request headers can be used when AWS_BEDROCK_ALLOW_GUARDRAIL_OVERRIDE is true:

  • No global guardrail configured: Setting is automatically true at startup, enabling per-request guardrails
  • Global guardrail configured: Setting defaults to false for security; set to true to allow overrides

This prevents users from bypassing configured safety controls while still allowing flexibility when no global policy exists.

Use HTTP headers to specify guardrail settings per request:

Header Purpose Valid Values
X-Amzn-Bedrock-GuardrailIdentifier Guardrail ID Your guardrail identifier
X-Amzn-Bedrock-GuardrailVersion Guardrail version Version number (e.g., 1)
X-Amzn-Bedrock-Trace Trace level disabled, enabled, enabled_full
X-Amzn-Bedrock-GuardrailStreamProcessingMode Guardrail assessment timing for streaming requests (stripped from non-streaming requests) sync, async
Example cURL Request
curl -X POST https://api.example.com/v1/chat/completions \
  -H "Authorization: Bearer sk-..." \
  -H "X-Amzn-Bedrock-GuardrailIdentifier: abc123def456" \
  -H "X-Amzn-Bedrock-GuardrailVersion: 1" \
  -H "X-Amzn-Bedrock-Trace: enabled" \
  -d '{"model": "anthropic.claude-sonnet-5", "messages": [...]}'

Request Body Configuration

The amazon-bedrock-guardrailConfig object in the request body is supported for OpenAI Chat Completions compatibility.

Compatibility Note

Only fields compatible with Bedrock Converse API are honored. The tagSuffix field is documented in AWS but not supported in this implementation.


Bedrock Session Storage

Requests with store=true on the Responses and Chat Completions APIs persist generations in Amazon Bedrock sessions. No environment variable is needed to enable this — it requires the Bedrock Session Storage IAM permissions.

Not available in every region

Amazon Bedrock session storage covers fewer regions than model inference. When the primary Bedrock region — the first entry of AWS_BEDROCK_REGIONS, which is where all sessions are created — does not provide it, store=true is ignored: the generation is still returned, and a warning is recorded in the request log stating that the session storage endpoint was unreachable or timed out and that session storage is offered in fewer regions than model inference. Retrieving a stored object then returns 404. A missing bedrock:CreateSession permission produces a distinct AccessDenied warning pointing at the IAM permissions instead.

Nothing fails and no request is lost, but stored responses and stored chat completions are simply unavailable. To rely on them, make the primary Bedrock region one that offers session storage — check the Amazon Bedrock session management endpoints for current coverage.

AWS_BEDROCK_SESSION_ENCRYPTION_KEY_ARN

Purpose : KMS key ARN encrypting the Amazon Bedrock sessions that back stored responses and stored chat completions (store=true)

Default : None — sessions are encrypted with the AWS-managed key

Validation : Checked at startup: must be a KMS key ARN (arn:<partition>:kms:<region>:<account-id>:key/<key-id>).

export AWS_BEDROCK_SESSION_ENCRYPTION_KEY_ARN=arn:aws:kms:us-east-1:123456789012:key/abcd-...

Shared Visibility Across Deployments

Stored responses and chat completions are namespaced by AWS account and region, not by stdapi.ai deployment. Multiple deployments sharing the same account and region can list, retrieve, and delete each other's stored objects. Use a dedicated AWS account per deployment when isolation matters, or accept this shared visibility as a deliberate trade-off.

Orphaned Session Cleanup

A session is created independently of the generation it will hold — before it for the Responses API, concurrently with it for Chat Completions — so a crash before the generation is written leaves an empty, orphaned session. Bedrock sessions have no TTL and persist until deleted, so periodically clean up stale sessions (aws bedrock-agent-runtime list-sessions plus delete-session, or an operator-managed lifecycle policy).

AWS_BEDROCK_BATCH_ROLE_ARN

Purpose : AWS IAM service role that Amazon Bedrock batch inference assumes to read a batch's requests and write its results — required to enable the Batch API and the Message Batches API

Type : String — an IAM role ARN

Default : None — the batch endpoints answer 503 (529 on the Anthropic-compatible routes) and the server reports the disabled feature in its startup log

Behavior : The server passes this role when it submits a batch; Amazon Bedrock then reads the requests and writes the results with it. The role must be able to read and write every bucket configured with AWS_S3_BUCKET and AWS_S3_REGIONAL_BUCKETS, under AWS_S3_BATCHES_PREFIX.

Validation : Checked at startup: must be an IAM role ARN (arn:<partition>:iam::<account-id>:role/<name>). Leaving it unset is reported as a startup warning, never a failure.

export AWS_BEDROCK_BATCH_ROLE_ARN=arn:aws:iam::123456789012:role/stdapi-ai-batch

The role and two IAM policies come first

The role's trust policy must allow bedrock.amazonaws.com to assume it, and the server's own role needs iam:PassRole on this ARN. See IAM Permissions for copyable policies.

AWS_BEDROCK_USER_ROLE_ARN

Purpose : Run each end user's model calls under an AWS IAM role session of their own, so AWS reports Amazon Bedrock model usage per end user in Cost Explorer and in the Cost and Usage Report

Type : String — an IAM role ARN

Default : None — every request runs under the server's own identity, and AWS reports all model usage under it

Behavior : The server opens one short-lived session of this role per end user, caches it, and signs that user's model invocations with it. The identity is the authenticated caller when authentication is enabled, otherwise the identifier the request declares (safety_identifier or user on the OpenAI-compatible APIs, metadata.user_id on the Anthropic Messages API). Only model invocations are covered — guardrail evaluations, video generation, speech, transcription and translation keep the server's identity. A session that cannot be opened fails the request rather than falling back to the server's identity.

Validation : Checked at startup: must be an IAM role ARN (arn:<partition>:iam::<account-id>:role/<name>). The server also tries to assume it at startup and reports a warning — not a failure — when it cannot.

export AWS_BEDROCK_USER_ROLE_ARN=arn:aws:iam::123456789012:role/stdapi-ai-end-user

The role and two IAM policies come first

The role's trust policy must allow this server's own role to call both sts:AssumeRole and sts:TagSession on it, and the server's role needs the same two actions on this role ARN. See IAM Permissions for copyable policies, including the model ARNs a cross-region inference profile requires.

AWS_BEDROCK_USER_ROLE_SESSION_DURATION

Purpose : Lifetime of a per-end-user role session, in seconds

Type : Integer — 900 to 3600

Default : 3600

Behavior : Sessions are cached per end user and reopened shortly before they expire, so a longer lifetime means fewer AWS STS calls. The upper bound is imposed by AWS: the server itself runs under an assumed role, and a role session obtained from another role session cannot last longer than one hour, whatever the role's maximum session duration.

export AWS_BEDROCK_USER_ROLE_SESSION_DURATION=1800

AWS_BEDROCK_USER_ROLE_TAG_KEY

Purpose : Session tag key carrying the end user identity on each per-end-user role session

Type : String, or null to send no session tag

Default : user

Behavior : Activate this key as a cost allocation tag — in the AWS Billing console, under Cost allocation tags filtered by type IAM principal — to group Bedrock costs by end user in Cost Explorer. The same tag is testable in IAM policies as aws:PrincipalTag/<key>, so the role can be restricted per user. With no tag, end users are still distinguished by their role session name in the Cost and Usage Report.

Validation : Checked at startup: 1 to 128 characters over letters, digits, spaces and _ . : / = + - @; keys beginning with aws: are reserved by AWS and rejected.

export AWS_BEDROCK_USER_ROLE_TAG_KEY=end-user

AWS_BEDROCK_USER_ROLE_REQUIRE_IDENTITY

Purpose : Reject a model request that identifies no end user, instead of running it under the server's own identity

Type : Boolean

Default : false — such requests run under the server's identity, and their usage is reported under it

Behavior : Enable it so that no model usage escapes per-user attribution: a request carrying neither an authenticated caller nor an end user identifier is answered 400. Clients that never send one stop working, so enable it only once every client identifies its user — and note that some APIs, audio transcription among them, have no end user field at all, so on those it takes an authenticated caller. A real-time speech-to-speech session keeps for its whole life the identity it opened with, so while this is enabled it is refused rather than attributed to the server. Requires AWS_BEDROCK_USER_ROLE_ARN.

export AWS_BEDROCK_USER_ROLE_REQUIRE_IDENTITY=true

Bedrock Service Tier and Performance Configuration

Amazon Bedrock service tiers and performance configurations allow you to optimize AI workload performance and cost trade-offs. Configure latency optimization and throughput priority for your inference requests.

AWS Documentation

For detailed information about service tiers, see:

Service Tiers

Service tiers help you match AI workload performance with cost by selecting the appropriate throughput and latency characteristics:

  • priority - Highest priority processing with guaranteed capacity and fastest response times. Best for latency-sensitive applications.
  • default - Standard processing with balanced performance and cost. Suitable for most production workloads.
  • flex - Cost-optimized processing with flexible scheduling. Best for batch jobs and non-time-sensitive workloads.

Performance Configuration

Performance configuration allows you to optimize for latency:

  • standard - Standard latency profile with balanced performance
  • optimized - Optimized for lowest possible latency

Per-Request Service Tier Configuration

Configure service tier and performance settings per request using HTTP headers. These headers are available on all Bedrock-based routes (Chat Completions, Embeddings, Images). Server-side per-model defaults can be set with DEFAULT_MODEL_SERVICE_TIERS.

Header Purpose Valid Values
X-Amzn-Bedrock-Service-Tier Service tier selection priority, default, flex
X-Amzn-Bedrock-PerformanceConfig-Latency Latency optimization standard, optimized

The tier header is subject to the override gate

When AWS_BEDROCK_ALLOW_SERVICE_TIER_OVERRIDE is false, X-Amzn-Bedrock-Service-Tier — like the service_tier request parameter — is ignored for any model that has a tier configured, by DEFAULT_MODEL_SERVICE_TIERS or by the alias the request names. A model with no configured tier honors the header in either case. The response's service_tier field keeps echoing the request's own value; usage and cost reporting record the tier that actually served the call.

Configured tiers, the header and this gate all apply to models served through the Bedrock Converse and InvokeModel APIs. On a Bedrock Mantle-served model, the request's own service_tier parameter is what applies — the header is not read, no configured tier is added, and the response reports the tier that model returns.

Example: Chat Completions with Priority Tier and Optimized Latency
curl -X POST https://api.example.com/v1/chat/completions \
  -H "Authorization: Bearer sk-..." \
  -H "X-Amzn-Bedrock-Service-Tier: priority" \
  -H "X-Amzn-Bedrock-PerformanceConfig-Latency: optimized" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "anthropic.claude-sonnet-5",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Example: Embeddings with Flex Tier for Batch Processing
curl -X POST https://api.example.com/v1/embeddings \
  -H "Authorization: Bearer sk-..." \
  -H "X-Amzn-Bedrock-Service-Tier: flex" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "amazon.nova-2-multimodal-embeddings-v1:0",
    "input": ["text 1", "text 2", "text 3"]
  }'
Example: Image Generation with Default Tier
curl -X POST https://api.example.com/v1/images/generations \
  -H "Authorization: Bearer sk-..." \
  -H "X-Amzn-Bedrock-Service-Tier: default" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "amazon.nova-canvas-v1:0",
    "prompt": "A serene mountain landscape"
  }'

When to Use Each Tier

Priority Tier:

  • Real-time customer-facing applications
  • Interactive chatbots and assistants
  • Applications requiring guaranteed low latency
  • Production workloads with strict SLAs

Default Tier:

  • Standard production workloads
  • General-purpose API usage
  • Applications with moderate latency requirements

Flex Tier:

  • Batch processing and bulk operations
  • Offline content generation
  • Data processing pipelines
  • Non-time-sensitive workloads
  • Cost-optimized inference at scale

Audio and Text-to-Speech

DEFAULT_TTS_MODEL

Purpose : Default text-to-speech model when not specified in requests

Default : amazon.polly-standard

Model Description Quality
amazon.polly-standard Standard Polly voices Classic quality
amazon.polly-neural Neural Polly voices Higher quality, more natural
amazon.polly-long-form Long-form content Optimized for long content
amazon.polly-generative Generative AI voices :material-sparkles: Latest technology
export DEFAULT_TTS_MODEL=amazon.polly-neural

DEFAULT_TTS_LANGUAGE

Purpose : Default language code for text-to-speech synthesis when using OpenAI voice names

Default : None (automatic language detection via Amazon Comprehend)

Behavior : When specified, this language is used instead of automatic detection. When not set, Amazon Comprehend detects the language automatically from the input text.

Valid Language Codes: Any Amazon Polly language code (e.g., en-US, fr-FR, es-ES, de-DE, ja-JP)

# Use English (US) for all TTS requests
export DEFAULT_TTS_LANGUAGE=en-US

# Use French for all TTS requests
export DEFAULT_TTS_LANGUAGE=fr-FR

Performance Benefits

Setting a default language improves performance by:

  • Faster responses: Skips language detection API call to Amazon Comprehend
  • Reduced costs: No Amazon Comprehend charges for language detection
  • Predictable voice selection: Always uses voices from the specified language

When to Use

Consider setting a default language when:

  • Your application primarily serves content in a single language
  • You want to optimize response times and reduce AWS service calls
  • You prefer predictable voice selection over automatic language matching

Interaction with Voice Selection

This setting only affects automatic language detection when using OpenAI voice names (like alloy, echo, nova). If you specify a Polly voice ID directly (like Joanna, Matthew), language detection is already skipped.


Realtime API

REALTIME_CLIENT_SECRET_KEY

Purpose : Secret the Realtime API's ephemeral client secrets (POST /v1/realtime/client_secrets) are signed with

Type : String (any value)

Default : None — a signing key is derived from the configured API key instead

Behavior : Ephemeral client secrets are stateless: nothing is stored server-side, so any instance behind a load balancer verifies a secret minted by any other, as long as they all sign with the same key. By default that key is derived from the deployment's own API key, which is already shared across every instance — so this setting has nothing to add on a deployment that already configures one.

Set it explicitly on a deployment that runs with **no API key configured at all**: without either one, each instance falls back to a random key generated **per process**, and a client secret minted by one instance then fails to verify on any other — the symptom is an ephemeral secret rejected intermittently on a multi-instance deployment. Any value works, as long as every instance shares it.
export REALTIME_CLIENT_SECRET_KEY=a-value-shared-by-every-instance

Changing it invalidates outstanding secrets

A secret minted with one key does not verify against another. Rotating this setting (or the API key it would otherwise derive from) invalidates every client secret minted before the change — they simply stop working once their bearer tries to open a session, the same as if they had expired.

REALTIME_ALLOW_SESSION_OVERRIDE

Purpose : Whether a client connecting with an ephemeral client secret may override the session configuration that secret carries

Type : Boolean

Default : true — the upstream behavior: the carried configuration is a default the client may change

Behavior : A minted secret carries a session configuration, and by default a client opening a session with it may name another model on the ?model= query string and replace any of that configuration with its own session.update — exactly as it can against the upstream API.

Set to `false` on a multi-tenant deployment, where the secret is the only thing constraining an untrusted browser or mobile client. The `model`, the `instructions` and `max_output_tokens` the secret was minted with are then final: connecting with a `?model=` naming a different model is refused before the session opens, and a `session.update` changing any of the three answers an `error` event. Everything else — voice, audio formats, turn detection, transcription — stays under the client's control.
export REALTIME_ALLOW_SESSION_OVERRIDE=false

Deprecated Settings

Deprecated and Ignored

TOKENS_ESTIMATION (default: false) and TOKENS_ESTIMATION_DEFAULT_ENCODING (default: None) are deprecated and ignored: tiktoken-based token estimation has been removed from the project. Token counts are now sourced directly from AWS billing data when available. Remove these variables from existing configurations.


Model Cache

stdapi.ai automatically discovers and caches available Bedrock models from configured regions. The cache is refreshed on-demand when expired, not via background tasks.

MODEL_CACHE_SECONDS

Purpose : Cache lifetime for the Bedrock models list before refresh

Type : Integer (seconds)

Default : 900 (15 minutes)

Behavior : When a request needs the model list (e.g., model lookup, /models endpoint) and the cache has expired, the server queries Amazon Bedrock to discover newly available models, check for model access changes, and update inference profile configurations. This cache also applies to application inference profile and prompt router information when users pass ARNs directly (if enabled via AWS_BEDROCK_ALLOW_APPLICATION_INFERENCE_PROFILE_ARN or AWS_BEDROCK_ALLOW_PROMPT_ROUTER_ARN)

# Default: 15 minutes
export MODEL_CACHE_SECONDS=900

# More frequent updates (5 minutes)
export MODEL_CACHE_SECONDS=300

# Less frequent updates (1 hour)
export MODEL_CACHE_SECONDS=3600

Lazy Refresh Behavior

The model cache uses lazy (on-demand) refresh, not background tasks:

  • Cache is refreshed only when a request needs it and the cache has expired
  • Common triggers: model lookup failures, /v1/models API calls, inference requests with unknown models
  • The first request after expiration experiences additional latency (typically 2-5 seconds) while the cache refreshes; the AWS calls (ListFoundationModels, GetFoundationModelAvailability, ListInferenceProfiles) run in parallel across regions, so the penalty scales with the slowest region rather than the number of regions
  • Subsequent requests use the fresh cache until it expires again

Tuning Recommendations

Interval Use Case Trade-offs
300 (5 min) Development, testing new models More frequent refresh latency, faster model discovery
900 (15 min) Production (default, balanced) Balanced refresh frequency and latency impact
3600 (1 hour) Stable production, cost optimization Rare refresh latency, slower model discovery

Lower cache lifetimes increase the frequency of the per-region discovery calls; very frequent refreshes in high-traffic deployments may approach API rate limits.


AI Response Timeout

AI_RESPONSE_TIMEOUT

Purpose : Maximum time in seconds to wait without receiving any data from an AI model

Type : Integer (seconds, must be greater than 0)

Default : 600 (10 minutes)

Behavior : Inactivity (per-read) timeout on the upstream model connection, applied to both streaming and non-streaming requests. The timer resets every time data is received, so it fires only when the model stalls for longer than this value — it does not bound the total duration of a response: a stream that keeps producing chunks can run well past it. On a non-streaming request, where the whole response arrives at once, it effectively bounds the wait for that single response. When it fires, the connection is closed and the request fails with a timeout error

# Default (10 minutes) - suitable for extended thinking models
export AI_RESPONSE_TIMEOUT=600

# Shorter timeout for standard models (2 minutes)
export AI_RESPONSE_TIMEOUT=120

# Longer timeout for very long documents or high reasoning budgets (15 minutes)
export AI_RESPONSE_TIMEOUT=900

When to Adjust

  • Increase if you see timeout errors with models that use extended thinking/reasoning, large document analysis, or high token budgets
  • Decrease to fail fast and free resources if your workload only uses standard models where long waits indicate a problem

Extended Thinking Models

Models with extended reasoning capabilities (such as Claude with thinking enabled or high reasoning_effort) may spend significant time generating internal reasoning steps before producing output. The default of 600 seconds accommodates these use cases. Standard models without extended thinking typically respond within 60 seconds.


Shutdown Drain

SHUTDOWN_DRAIN_TIMEOUT

Purpose : Maximum time in seconds the server waits for background work to finish after it has been asked to stop

Type : Number (seconds, 0 or greater)

Default : 10

Behavior : Some work is deliberately started outside the request that asked for it, so the caller is answered without waiting for it: temporary file cleanups, vector store file indexing, and the release of live audio sessions. On a stop signal the server waits up to this long for that work to finish, then cancels whatever is still running. The wait is a single deadline shared by all of it, not a budget per item, and a server with nothing outstanding stops immediately

# Default: comfortably inside a 30-second container stop timeout
export SHUTDOWN_DRAIN_TIMEOUT=10

# Longer wait, with the container stop timeout raised to match
export SHUTDOWN_DRAIN_TIMEOUT=20

# No wait: cancel background work immediately and stop as fast as possible
export SHUTDOWN_DRAIN_TIMEOUT=0

Best effort, not a delivery guarantee

A container runtime sends SIGKILL a fixed delay after the stop signal — 30 seconds by default on Amazon ECS — so the server can be killed before the wait ends, and a deployment may run under an orchestrator that stops it sooner still. Keep this value comfortably below your container stop timeout, and raise it only together with that timeout. Never rely on this wait for anything whose completion matters: retry the operation instead.

When Work Is Lost

Anything cancelled at the deadline is counted in the server's stop log event, which is then emitted at warning level with one count per kind of work. Deployments that see those counts regularly are stopping the server faster than its work can finish: raise this value and the container stop timeout together, or reduce what each request defers.


Default Model Parameters

Configure default inference parameters applied automatically to specific models.

What You Can Do

  • Set consistent temperature/creativity levels per model
  • Enable provider-specific features (e.g., Anthropic beta features)
  • Configure default token limits for cost control
  • Apply model-specific stop sequences

Parameter Precedence

Request parameters always take precedence over defaults.

DEFAULT_MODEL_PARAMS

Purpose : Per-model default parameters

Format : JSON object with model IDs as keys

Supported Parameters:

Parameter Type Range Description
temperature Float ≥ 0 Sampling temperature
top_p Float ≥ 0 Nucleus sampling
max_tokens Integer ≥ 1 Maximum response tokens
stop_sequences String/Array - Stop generation tokens
Provider-specific Various - e.g., anthropic_beta

Only the outer JSON shape (an object of per-model objects) is validated at startup. The parameter values above are validated lazily, the first time a model with configured defaults is used: a wrong type, or a value below the lower bounds shown in the table, fails that request with HTTP 400. The numeric ceilings (for example the usual top_p maximum of 1.0) are enforced by Amazon Bedrock and the target model.

Configuration Examples

Basic Parameters:

export DEFAULT_MODEL_PARAMS='{
  "amazon.nova-micro-v1:0": {
    "temperature": 0.3,
    "max_tokens": 800
  }
}'

Provider-Specific Features:

export DEFAULT_MODEL_PARAMS='{
  "anthropic.claude-sonnet-5": {
    "anthropic_beta": ["Interleaved-thinking-2025-05-14"]
  }
}'

Multiple Models:

export DEFAULT_MODEL_PARAMS='{
  "amazon.nova-micro-v1:0": {
    "temperature": 0.3,
    "max_tokens": 500
  },
  "amazon.nova-lite-v1:0": {
    "temperature": 0.7,
    "max_tokens": 2000
  },
  "anthropic.claude-sonnet-5": {
    "temperature": 0.5,
    "top_p": 0.9,
    "anthropic_beta": ["Interleaved-thinking-2025-05-14"]
  }
}'

Advanced Configuration:

export DEFAULT_MODEL_PARAMS='{
  "amazon.nova-pro-v1:0": {
    "temperature": 0.7,
    "top_p": 0.95,
    "max_tokens": 4096,
    "stop_sequences": ["Human:", "Assistant:"]
  }
}'

Parameter Merging

graph LR
    A[Default Parameters] --> B[Merged Config]
    C[Request Parameters] --> B
    B --> D[Final Configuration]
  1. Default parameters are applied first (from DEFAULT_MODEL_PARAMS)
  2. Request parameters override defaults if both are specified
  3. Provider-specific fields are forwarded to Bedrock as additional model request fields
  4. Unsupported fields reach Bedrock as-is, and a field the model rejects surfaces as a ValidationException returned to the client as HTTP 400. Three cases are handled before that: anthropic_beta flags are filtered individually against an allowlist (see ANTHROPIC_BETA_FILTER); a system prompt sent to a model that does not support one is dropped when DROP_UNSUPPORTED_SYSTEM_PROMPT is enabled (the default); and Amazon Nova 2 drops max_tokens when reasoning effort is high, logging a warning

Default Model Service Tiers

Configure default service tiers applied automatically to specific Bedrock models.

What You Can Do

  • Set cost-efficient tiers for batch and agentic workloads by default
  • Configure priority tiers for latency-sensitive models
  • Optimize compute costs without modifying client requests

Available Service Tiers

Tier Description
default Standard compute tier (default)
flex Flexible compute tier for cost optimization
priority Priority compute tier for lower latency
reserved Reserved capacity for dedicated resources (requires AWS contract)

When to Use Each Tier

  • Default: Everyday AI tasks like content generation and text analysis
  • Flex: Cost-sensitive workloads like model evaluations, summarization, and agentic workflows
  • Priority: Mission-critical applications requiring lowest latency
  • Reserved: Predictable workloads needing 99.5% uptime guarantee (requires AWS contact)

Model Support

Not all models support all service tiers. Check the official AWS documentation for each model's supported tiers.

Examples:

  • amazon.nova-pro-v1:0 supports: default, flex, priority (not reserved)
  • amazon.nova-premier-v1:0 (legacy) supports: default, flex, priority, reserved

Tier Precedence

Explicit request parameters take precedence over the tier configured for the model, unless AWS_BEDROCK_ALLOW_SERVICE_TIER_OVERRIDE is disabled.

DEFAULT_MODEL_SERVICE_TIERS

Purpose : Per-model default service tier

Format : JSON object with model IDs as keys and tier string as value

Default : {}

Supported Values:

Value Description
default Standard compute (Bedrock default)
flex Cost-optimized flexible compute
priority Lower-latency priority compute
reserved Dedicated reserved capacity (requires AWS contract)

AWS_BEDROCK_ALLOW_SERVICE_TIER_OVERRIDE

Purpose : Control whether clients can select the service tier at request level

Default : true (clients may select a tier)

:octicons-cash-24: Cost Consideration : Service tiers are billed at different rates. Set to false on a shared deployment to pin every model to the tier you configured, so a client cannot move its traffic to a more expensive tier. A model with no configured tier still honors the request in either case.

Scope : Applies to models served through the Bedrock Converse and InvokeModel APIs. A Bedrock Mantle-served model carries no configured tier, so its requests always run on the tier they name.

export AWS_BEDROCK_ALLOW_SERVICE_TIER_OVERRIDE=false

Configuration Examples

Single Model:

export DEFAULT_MODEL_SERVICE_TIERS='{
  "amazon.nova-pro-v1:0": "flex"
}'

Multiple Models:

export DEFAULT_MODEL_SERVICE_TIERS='{
  "amazon.nova-pro-v1:0": "flex",
  "amazon.nova-premier-v1:0": "priority"
}'

Service Tier Merging

For models served through the Bedrock Converse and InvokeModel APIs:

  1. Explicit request parameter takes highest priority
  2. HTTP header (X-Amzn-Bedrock-Service-Tier, see Per-Request Service Tier Configuration) overrides defaults
  3. Tier configured on the requested alias (see Model Aliases) applies if the request sets none
  4. Default from DEFAULT_MODEL_SERVICE_TIERS applies if neither does
  5. No service tier passed to Bedrock if unset

Steps 1 and 2 are skipped when AWS_BEDROCK_ALLOW_SERVICE_TIER_OVERRIDE is false and a tier is configured.

On a Bedrock Mantle-served model, only step 1 applies: the request's own service_tier is forwarded as sent, and neither the header, nor an alias' tier, nor DEFAULT_MODEL_SERVICE_TIERS takes part.


Model Aliases

Configure custom aliases to map user-friendly model names to actual model IDs. This enables OpenAI API compatibility and simplifies model references.

What You Can Do

  • Create custom aliases for frequently used models
  • Enable OpenAI-compatible model names by default
  • Simplify model ID references in API requests
  • Seamlessly migrate between model versions

Default Aliases

stdapi.ai includes default aliases for OpenAI compatibility:

  • tts-1amazon.polly-standard
  • tts-1-hdamazon.polly-neural
  • whisper-1amazon.transcribe

stdapi.ai also supports dynamic model name aliases matching official provider APIs (OpenAI, Anthropic). You can use model names from provider documentation (e.g., claude-sonnet-5, gpt-oss-20b) which are automatically resolved to their corresponding Amazon Bedrock model identifiers.

MODEL_ALIASES

Purpose : Map alias names to actual model IDs or ARNs

Format : JSON object with alias names as keys, and as values either a model ID or ARN, or an object carrying that model plus the configuration to apply to it

Default : {} (empty, uses built-in defaults only)

Advanced Routing with ARNs

Model aliases can also reference ARNs for Application Inference Profiles or Prompt Routers, enabling advanced routing strategies through friendly alias names. See Using Inference Profile and Prompt Router ARNs for more details.

Aliases That Carry Configuration

An alias may map to an object instead of a model name. Every request naming that alias then gets the configuration attached to it, so one deployment can publish the same model under several names with different tiers, safeguards or defaults.

Every field below is designed around an Amazon Bedrock model call. An alias pointing at a model served by another AWS service — Amazon Polly, Amazon Transcribe, Amazon Comprehend — still resolves the name, and two fields keep working there: extra_params applies wherever the route accepts model parameters, and guardrail_id is enforced on the routes that check content through Amazon Bedrock Guardrails (speech input, transcripts) — a billed guardrail evaluation. service_tier and metadata configure the Bedrock call itself and are ignored on those services.

Field Purpose
model Required. Model ID or ARN the alias resolves to
service_tier Service tier for requests naming the alias, on a model served through the Bedrock Converse or InvokeModel APIs — see Default Model Service Tiers
guardrail_id ID of an Amazon Bedrock Guardrail to apply, requires guardrail_version; guardrail_identifier is accepted as the same field
guardrail_version Version of that guardrail
guardrail_trace Guardrail trace level: disabled, enabled or enabled_full
metadata Key-value metadata attached to the model call, for audit reporting — it reaches Amazon Bedrock model invocation logs, which you enable and deliver yourself, and nothing else: it is not a cost allocation tag, see AWS Cost Attribution
extra_params Model parameters, in the format of DEFAULT_MODEL_PARAMS
export MODEL_ALIASES='{
  "support-assistant": {
    "model": "amazon.nova-lite-v1:0",
    "service_tier": "flex",
    "guardrail_id": "abc123def456",
    "guardrail_version": "1",
    "metadata": {"team": "support"},
    "extra_params": {"temperature": 0.2}
  }
}'

Precedence

Each field resolves in one order: the request, then the alias, then the server-wide setting for that field. A field the alias leaves unset falls through to the server-wide value, and a client that sends nothing gets the alias' configuration.

The two settings that decide whether a request may override an administrator's value apply to the alias layer as well:

Startup Validation

An alias object is validated when the server starts: an unknown field, a missing model, a guardrail ID without its version, or an out-of-range extra_params value stops startup with an error naming the alias. A typo never becomes a silently ignored setting.

An alias whose guardrail_id targets a model served through Bedrock Mantle also stops startup: Amazon Bedrock Guardrails do not apply to those models, and serving them unfiltered while a guardrail is configured would be a silent gap. Point the alias at another model, or — when the model is also available on the classic endpoint — remove it from AWS_BEDROCK_MANTLE_PREFERRED_MODELS so it is served where guardrails apply.

Scope on Bedrock Mantle models

On a Bedrock Mantle-served model, guardrail_id is rejected at startup as above, and service_tier, metadata and extra_params — like the server-wide DEFAULT_MODEL_SERVICE_TIERS and DEFAULT_MODEL_PARAMS — do not apply. Such a request runs on the tier it names itself, and on that model's default tier when it names none.

Configuration Examples

Basic Alias:

export MODEL_ALIASES='{
  "my-tts": "amazon.polly-neural",
  "my-stt": "amazon.transcribe"
}'

Override Default Aliases:

# Override the default tts-1 mapping
export MODEL_ALIASES='{
  "tts-1": "amazon.polly-generative"
}'

Multiple Custom Aliases:

export MODEL_ALIASES='{
  "fast-model": "amazon.nova-micro-v1:0",
  "balanced-model": "amazon.nova-lite-v1:0",
  "quality-model": "amazon.nova-pro-v1:0",
  "claude": "anthropic.claude-sonnet-5"
}'

Map OpenAI Models to Bedrock:

# Make OpenAI model names work with Amazon Bedrock models
export MODEL_ALIASES='{
  "gpt-5": "anthropic.claude-sonnet-5",
  "gpt-4o": "anthropic.claude-sonnet-5",
  "gpt-4o-mini": "anthropic.claude-haiku-4-5-20251001-v1:0",
  "dall-e-3": "amazon.nova-canvas-v1:0",
  "dall-e-2": "stability.stable-image-ultra-v1:1"
}'

Override Deprecated Models:

# Redirect deprecated model IDs to their newer replacements
export MODEL_ALIASES='{
  "amazon.titan-image-generator-v1": "amazon.nova-canvas-v1:0",
  "amazon.titan-text-express-v1": "amazon.nova-lite-v1:0",
  "anthropic.claude-3-5-sonnet-20240620-v1:0": "anthropic.claude-sonnet-5",
  "stability.stable-image-ultra-v1:0": "stability.stable-image-ultra-v1:1"
}'

Advanced Routing with ARNs:

# Map friendly names to Application Inference Profiles or Prompt Routers
export MODEL_ALIASES='{
  "my-router": "arn:aws:bedrock:us-east-1:123456789012:default-prompt-router/cost-optimizer",
  "my-profile": "arn:aws:bedrock:us-east-1:123456789012:application-inference-profile/abc123xyz",
}'

Using Aliases in API Requests

Once configured, aliases can be used anywhere a model ID is expected:

# Using the default tts-1 alias
curl https://api.example.com/v1/audio/speech \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Hello world",
    "voice": "alloy"
  }'

# Using a custom alias
curl https://api.example.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "fast-model",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Alias Resolution

graph LR
    A[API Request] --> B{Alias Exists?}
    B -->|Yes| C[Resolve to Model ID]
    B -->|No| D[Use as Model ID]
    C --> E[Model Validation]
    D --> E
    E --> F[Execute Request]
  1. User-configured aliases override default aliases
  2. Default aliases apply if not overridden
  3. Non-aliased names pass through unchanged
  4. Resolved model ID is validated and used for the request

System Prompt Handling

Control how system prompts are handled for models that don't support them.

DROP_UNSUPPORTED_SYSTEM_PROMPT

Purpose : Control system prompt behavior for models that don't support system prompts

Type : Boolean

Default : true

# Default: silently drop system prompts for unsupported models
export DROP_UNSUPPORTED_SYSTEM_PROMPT=true

# Strict mode: return error when system prompt is used with unsupported model
export DROP_UNSUPPORTED_SYSTEM_PROMPT=false

Models Without System Prompt Support

Some Bedrock models don't support system prompts, including:

  • mistral.mistral-7b-instruct-v0:2
  • mistral.mixtral-8x7b-instruct-v0:1
  • Other older or specialized models

Use Cases

Enable (true, default) for:

  • Backward compatibility - Existing applications continue working
  • Model flexibility - Switch between models without code changes
  • Graceful degradation - System prompts are ignored instead of failing
  • Global system prompts - Applications that set system prompts globally for all models work seamlessly

Disable (false) for:

  • Strict validation - Catch configuration errors early
  • Debugging - Identify when system prompts aren't being used
  • Security requirements - Ensure system prompts are always applied

Anthropic Beta Flag Filtering

Anthropic-compatible clients like Claude Code send anthropic-beta headers with experimental beta flags. Many of these flags (such as files-api-2025-04-14, prompt-caching-2024-07-31) are not supported by Amazon Bedrock and cause ValidationException errors (HTTP 400).

stdapi.ai automatically filters out unsupported flags while preserving supported ones, so clients work without any special configuration. Previously, the workaround was to set CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1 on the client side, but this also disabled Bedrock-supported flags like Interleaved-thinking-2025-05-14 and token-efficient-tools-2025-02-19, degrading capabilities. This workaround is no longer needed.

Filtering is controlled by two settings: ANTHROPIC_BETA_FILTER to enable or disable it, and ANTHROPIC_BETA_ALLOWLIST to extend the built-in set of allowed flags.

ANTHROPIC_BETA_FILTER

Purpose : Enable or disable filtering of unsupported anthropic_beta flags for Anthropic Claude models

Type : Boolean

Default : true

Behavior : When enabled, anthropic_beta flags not in the allowlist are silently removed from requests before they reach Bedrock. A warning is logged when flags are filtered. When disabled, all flags are passed through to Bedrock as-is

# Enabled (default) - filter unsupported flags automatically
# No environment variable needed

# Disable filtering entirely (pass all flags through to Bedrock)
export ANTHROPIC_BETA_FILTER=false

When to Disable

Set to false only when:

  • Testing - You want to verify Bedrock behavior with specific flags directly
  • Custom setups - You manage flag compatibility at the client level

ANTHROPIC_BETA_ALLOWLIST

Purpose : Add extra anthropic_beta flags to the built-in set of Bedrock-supported flags

Format : Comma-separated string of additional beta flag names

Default : Empty (only the built-in Bedrock defaults are used)

Behavior : The flags specified here are merged with the built-in set of Bedrock-supported flags. You only need to specify extra flags beyond the defaults (e.g., newly added Bedrock flags). Only effective when ANTHROPIC_BETA_FILTER is true

# Use built-in defaults only (recommended) - no environment variable needed

# Add newly supported Bedrock flags without waiting for a stdapi.ai update
export ANTHROPIC_BETA_ALLOWLIST='new-feature-2026-03-01,another-flag-2026-04-01'

Built-in Allowed Flags:

Flag Feature
computer-use-2024-10-22 Computer use (Claude 3.5)
computer-use-2025-01-24 Computer use (Claude 3.7)
computer-use-2025-11-24 Computer use (Claude 4.5/4.6)
token-efficient-tools-2025-02-19 Token efficient tools
Interleaved-thinking-2025-05-14 Interleaved thinking
output-128k-2025-02-19 128K output
dev-full-thinking-2025-05-14 Raw thinking dev mode
context-1m-2025-08-07 1M context
context-management-2025-06-27 Context management (memory)
effort-2025-11-24 Effort control
tool-search-tool-2025-10-19 Tool search
tool-examples-2025-10-29 Tool use examples

Use Cases

Filtering enabled (default) for:

  • Claude Code via Bedrock - Clients work without CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1
  • Production stability - Prevent unsupported flags from causing request failures
  • Drop-in compatibility - Clients configured for direct Anthropic API work through stdapi.ai without changes

EXTRA_MODEL_PARAMS_DENYLIST

Purpose : Add extra parameter names to strip from the "extra model parameters" passthrough (any undeclared top-level JSON field on a chat or non-chat route, forwarded to Bedrock as a provider-specific inference field)

Format : Comma-separated string of additional parameter names

Default : Empty (only the built-in denylist is used)

Behavior : The names specified here are merged with the built-in denylist of LiteLLM client-control parameters (such as drop_params, api_key, custom_llm_provider) that some OpenAI-SDK-based clients leak into extra_body and that are never legitimate Bedrock model parameters — for example RAGFlow hardcodes extra_body={"drop_params": True} on every embeddings call, which previously reached Bedrock as an unrecognized inference field and failed with ValidationException. Every other extra parameter keeps being forwarded as before. Only effective when EXTRA_MODEL_PARAMS_DROP_ALL is false

# Use the built-in denylist only (recommended) - no environment variable needed

# Also strip a project-specific control field some client leaks into requests
export EXTRA_MODEL_PARAMS_DENYLIST='x_internal_debug_flag,x_proxy_trace_id'

EXTRA_MODEL_PARAMS_DROP_ALL

Purpose : Disable the "extra model parameters" passthrough entirely

Type : Boolean

Default : false

Behavior : When enabled, no undeclared request field is ever forwarded to Bedrock as a provider-specific inference parameter, on every route that supports the passthrough (chat completions/responses/messages, and embeddings/images/audio/rerank/etc.). This overrides EXTRA_MODEL_PARAMS_DENYLIST: with drop-all enabled, denylist filtering no longer matters because nothing is forwarded. Per-model defaults configured through DEFAULT_MODEL_PARAMS are unaffected — only request-supplied extras are dropped

# Keep the passthrough (default) - no environment variable needed

# Lock the deployment down to only declared API fields
export EXTRA_MODEL_PARAMS_DROP_ALL=true

When to Enable

Set to true only when you need to guarantee that no undeclared client field ever reaches Bedrock, for example a strict multi-tenant deployment where provider-specific knobs must go through an explicit allowlisted mechanism instead of the passthrough.


Image Generation

IMAGE_GENERATION_MODEL

Purpose : Default Bedrock image model ID used when the image_generation integrated tool is invoked via the Responses API. The tool intercepts requests from any text model, generates the image against this Bedrock image model, and returns an image_generation_call output item.

Type : String (Bedrock image model ID)

Default : None — the tool returns HTTP 400 if no model is configured and the request does not specify one

Behavior : The tool definition in the request may include a model field to override this default per call. Priority: request model field > this env var. Any available Bedrock image generation model can be used — for example amazon.nova-canvas-v1:0, amazon.titan-image-generator-v2:0, or the Stability AI Stable Image / Stable Diffusion family. Legacy models (such as amazon.titan-image-generator-v1 and stability.stable-diffusion-xl-v1) are hidden unless AWS_BEDROCK_LEGACY is enabled. Use the Search Models API to list the image models available in your deployment.

export IMAGE_GENERATION_MODEL='amazon.nova-canvas-v1:0'

With this set, any text model can generate images via the Responses API:

curl -X POST "$BASE/v1/responses" \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "amazon.nova-micro-v1:0",
    "input": "Generate a sunset over the ocean.",
    "tools": [{"type": "image_generation"}],
    "tool_choice": "required"
  }'

Using Inference Profile and Prompt Router ARNs

stdapi.ai supports passing ARNs directly as model IDs in API requests, enabling advanced routing capabilities beyond standard model selection.

Simplify ARNs with Model Aliases

Instead of using long ARNs directly in API requests, you can create Model Aliases that map friendly names to ARNs. This provides shorter, easier-to-use naming for your API users.

Overview

Instead of using standard model IDs like anthropic.claude-sonnet-5, you can pass ARNs that reference:

  • Cross-Region Inference Profiles - AWS-managed multi-region routing
  • Application Inference Profiles - Your custom routing configurations
  • Prompt Routers - Intelligent dynamic model selection

Automatic Cross-Region Routing

stdapi.ai automatically handles cross-region routing by default. When you use standard model IDs, the application automatically selects and uses the optimal AWS-managed cross-region inference profile based on your configured AWS_BEDROCK_REGIONS.

You typically do not need to manually pass cross-region inference profile ARNs. The automatic selection handles routing across your configured regions for best availability and latency.

Manual ARN passing is primarily useful for:

  • Application inference profiles - Your custom routing configurations
  • Prompt routers - Intelligent cost optimization and dynamic model selection
  • Rare cases - When you need to override automatic cross-region profile selection

Enabling ARN Support

By default, users can only pass standard model IDs. To allow ARN usage, enable the appropriate settings:

# Allow cross-region inference profile ARNs
export AWS_BEDROCK_ALLOW_CROSS_REGION_INFERENCE_PROFILE_ARN=true

# Allow application inference profile ARNs
export AWS_BEDROCK_ALLOW_APPLICATION_INFERENCE_PROFILE_ARN=true

# Allow prompt router ARNs
export AWS_BEDROCK_ALLOW_PROMPT_ROUTER_ARN=true

Security Consideration

These settings are disabled by default. Only enable them when you want to give users explicit control over ARN-based routing. For centralized server-controlled routing, use AWS_BEDROCK_MODEL_ARN_MAPPING instead.

Using ARNs in API Requests

Once enabled, users can pass ARNs directly in the model parameter:

Cross-Region Inference Profile Example:

curl -X POST https://api.example.com/v1/chat/completions \
  -H "Authorization: Bearer sk-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arn:aws:bedrock:us-east-1:123456789012:inference-profile/us.anthropic.claude-sonnet-5",
    "messages": [
      {"role": "user", "content": "Hello!"}
    ]
  }'

Application Inference Profile Example:

curl -X POST https://api.example.com/v1/chat/completions \
  -H "Authorization: Bearer sk-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arn:aws:bedrock:us-east-1:123456789012:application-inference-profile/my-custom-profile",
    "messages": [
      {"role": "user", "content": "Hello!"}
    ]
  }'

Prompt Router Example:

curl -X POST https://api.example.com/v1/chat/completions \
  -H "Authorization: Bearer sk-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "arn:aws:bedrock:us-east-1:123456789012:default-prompt-router/my-router",
    "messages": [
      {"role": "user", "content": "Hello!"}
    ]
  }'

Use Case Comparison

Approach Best For Configuration
Standard Model IDs Most common use case, simple routing No special configuration needed
Server-Side ARN Mapping Centralized control, transparent to clients AWS_BEDROCK_MODEL_ARN_MAPPING
Client-Side ARN Passing User-controlled routing, advanced use cases Enable AWS_BEDROCK_ALLOW_*_ARN settings

Best Practices

Recommended Approach

For most deployments, use server-side ARN mapping (AWS_BEDROCK_MODEL_ARN_MAPPING):

  • Centralized control over routing behavior
  • Transparent to API clients
  • Easy to change routing without modifying client code
  • Better security (server controls which ARNs are used)

When to Allow Client-Side ARNs

Enable AWS_BEDROCK_ALLOW_*_ARN settings when:

  • Clients need fine-grained control over routing
  • Different clients require different routing strategies
  • Advanced users managing their own inference profiles
  • Testing and comparing different routing configurations

Security and Governance

When enabling client-side ARN passing:

  • Clients can bypass server-configured routing
  • Monitor usage to prevent unexpected costs
  • Ensure appropriate IAM permissions are in place
  • Track ARN usage through logs and monitoring

Required IAM Permissions

When using ARN-based routing, ensure your IAM role/user has the appropriate permissions:

{
  "Sid": "BedrockARNRouting",
  "Effect": "Allow",
  "Action": [
    "bedrock:GetInferenceProfile",
    "bedrock:GetPromptRouter"
  ],
  "Resource": "*"
}

See the IAM Permissions page for complete policy examples.


Next Steps