Configuration Guide¶
stdapi.ai is configured entirely through environment variables, which are read once at startup and cannot be changed without restarting the service. This guide explains each setting category with practical examples to help you configure the service correctly.
What you can configure:
- AWS regions - Access models across multiple regions for availability and model selection
- Data sovereignty - Control which AWS regions are used for compliance (GDPR, HIPAA, etc.)
- Storage - S3 buckets for file operations, regional buckets for multi-region deployments
- Authentication - API keys via SSM or Secrets Manager for secure access control
- Observability - Logging levels, OpenTelemetry, request/response debugging
- Security - CORS, proxy headers, trusted hosts for production deployments
- Performance - Caching, model overrides, S3 acceleration
- TLS / SSL - End-to-end encryption using Granian environment variables
Secure Defaults on Startup
Every setting below has a default, so stdapi.ai starts with none of them set: it uses the AWS credentials it runs with, detects your current AWS region, and discovers the available Bedrock models.
Prerequisites
Before configuring stdapi.ai, ensure you have:
- AWS Account with access to Amazon Bedrock
- AWS Credentials configured via environment variables, AWS CLI, or IAM role (for EC2/ECS/Lambda deployments)
- IAM Permissions to access required AWS services (see the IAM Permissions guide)
- S3 Bucket (optional, but recommended for production use with file operations)
Container Runtime
Both the AWS Marketplace and community Docker images run using Granian, a high-performance Python ASGI server. In addition to the stdapi.ai-specific configuration variables documented below, you can also use Granian environment variables to configure the server runtime (e.g., GRANIAN_PORT, GRANIAN_WORKERS, GRANIAN_THREADS, etc.).
The images listen on IPv4 only (GRANIAN_HOST=0.0.0.0). Set GRANIAN_HOST=:: to bind a dual-stack socket answering both IPv4 and IPv6 clients. This is needed wherever a client may resolve the server to an IPv6 address — in particular with ECS service discovery, which publishes an AAAA record for every task in an IPv6-enabled subnet, and some clients (Node.js among them) try that address first and fail with ECONNREFUSED against an IPv4-only listener. The official Terraform module sets it for you when the VPC has IPv6 enabled.
Quick Start¶
For production deployments, configure these essential settings:
Minimal Production Setup¶
Single-region deployment with file storage only.
# S3 bucket for file storage (must be in same region as your server)
export AWS_S3_BUCKET=my-stdapi-bucket
# AWS_BEDROCK_REGIONS is optional - will auto-detect your current AWS region if not specified
Production with Authentication¶
Adds secure API key authentication via AWS Systems Manager.
# S3 bucket for file storage (must be in same region as your server)
export AWS_S3_BUCKET=my-stdapi-bucket
# Secure API authentication (recommended: SSM Parameter Store)
export API_KEY_SSM_PARAMETER=/stdapi/prod/api-key
# AWS_BEDROCK_REGIONS is optional - will auto-detect your current AWS region if not specified
Full Production Setup (All Features Enabled)¶
Multi-region deployment with all AWS AI services, observability, and security features.
# Core AWS configuration - host server in first region
export AWS_BEDROCK_REGIONS=us-east-1,us-west-2,eu-west-1
# S3 bucket for file storage (must be in us-east-1, your first/primary region)
export AWS_S3_BUCKET=my-stdapi-us-east-1-bucket
# Optional: Transcribe S3 bucket (defaults to AWS_S3_BUCKET if not specified)
# Only set this if you need a separate bucket or if transcribe is in a different region
# export AWS_TRANSCRIBE_S3_BUCKET=my-stdapi-transcribe-us-east-1
# Optional: Regional buckets for async/batch inference in other regions
export AWS_S3_REGIONAL_BUCKETS='{"us-west-2": "my-stdapi-us-west-2-bucket", "eu-west-1": "my-stdapi-eu-west-1-bucket"}'
# AWS AI services regions (optional - when unset, every AWS_BEDROCK_REGIONS entry is a
# candidate with automatic failover; set one to pin the service to a single region)
export AWS_POLLY_REGION=us-east-1 # Text-to-speech
export AWS_TRANSCRIBE_REGION=us-east-1 # Speech-to-text (audio transcription)
export AWS_COMPREHEND_REGION=us-east-1 # Language detection & moderation
export AWS_TRANSLATE_REGION=us-east-1 # Text translation
# Authentication
export API_KEY_SSM_PARAMETER=/stdapi/prod/api-key
# Logging
export LOG_LEVEL=warning
export LOG_CLIENT_IP=true
# Optional: OpenTelemetry observability (AWS X-Ray integration)
# export OTEL_ENABLED=true
# export OTEL_SERVICE_NAME=stdapi-production
# export OTEL_SAMPLE_RATE=0.1
# Production security settings (when behind AWS ALB/CloudFront)
export ENABLE_PROXY_HEADERS=true
# Note: TRUSTED_HOSTS not recommended with AWS ALB - use ALB host-based routing instead
# Only use TRUSTED_HOSTS if you cannot configure host validation at the load balancer level
# Optional: CORS for browser-based web applications
# export CORS_ALLOW_ORIGINS='["https://app.example.com"]'
Development Setup¶
Local development configuration with API documentation and debug logging enabled.
# Minimal configuration for local development
export AWS_S3_BUCKET=my-stdapi-dev-bucket
# Enable API documentation
export ENABLE_DOCS=true
export ENABLE_REDOC=true
# Full request/response logging for debugging
export LOG_LEVEL=info
export LOG_REQUEST_PARAMS=true
# AWS_BEDROCK_REGIONS is optional - will auto-detect your current AWS region if not specified
S3 Bucket Required for Certain Features
Without an S3 bucket configured, some features will be disabled (such as image output as URL, audio transcription). See the relevant API documentation for feature requirements.
All Other Settings Are Optional
The configurations above are sufficient for most production deployments. All other settings can be configured as needed for your specific use case.
What to Set¶
Most deployments set two variables. The official Terraform module sets both for you. Everything else has a working default; each tier below tells you when to leave that default behind.
Two flags mark the rows to read in full before setting them: security changes who can reach the API or what a caller may do, cost adds or changes an AWS bill line.
Tier 1 — Set these (2)¶
| Variable | What it does | Set it when | Flags |
|---|---|---|---|
AWS_S3_BUCKET | Primary S3 bucket for file storage; must be in the first region of AWS_BEDROCK_REGIONS | Always, unless you never call a file, image, video, audio or batch endpoint | cost |
AWS_BEDROCK_REGIONS | Which regions serve models; the first is where the server should be hosted | You want models from more than one region, or a region other than the one the server runs in |
Tier 2 — Common (about 20)¶
| Variable | What it does | Set it when | Flags |
|---|---|---|---|
API_KEY_SSM_PARAMETER | Client API key read from SSM Parameter Store | Any deployment reachable by more than you | security |
AUTHENTICATION_MODE | Which credentials the gateway accepts (API key, Cognito token, or both) | You put an identity provider in front of the gateway | security |
AWS_BEDROCK_USER_ROLE_ARN | Per-end-user IAM role the gateway assumes for Bedrock calls | You need per-user attribution or per-user isolation in CloudTrail and billing | security |
AWS_S3_VECTORS_BUCKET | S3 vector bucket backing the Vector Stores API; unset disables it | You use file search or RAG | cost |
AWS_SQS_VECTOR_STORE_QUEUE_URL | SQS queue that makes vector store indexing outlive the server running it | Vector store indexing must survive a restart or a scale-in | cost |
AWS_DYNAMODB_TABLE | DynamoDB table holding the records a deployment's instances share | You run more than one instance and want conversations, sessions and caches shared | cost |
AWS_S3_REGIONAL_BUCKETS | Region-specific buckets for Bedrock async and batch inference | You run batch or async inference in more than one region | cost |
AWS_TRANSCRIBE_S3_BUCKET | Bucket for temporary transcription files | Transcribe runs in a region other than the one holding AWS_S3_BUCKET | cost |
AWS_BEDROCK_MANTLE_ENABLED | Exposes the models served by the Bedrock Mantle endpoint alongside the classic Bedrock Converse catalogue | You want the Mantle-served models off — they are on by default | cost |
AWS_BEDROCK_MARKETPLACE_ENDPOINTS_ENABLED | Exposes AWS Marketplace model endpoints | You serve a Marketplace model | cost |
AWS_BEDROCK_GUARDRAIL_IDENTIFIER | Applies a Bedrock guardrail to every eligible route | You must filter prompts or responses centrally | security cost |
DEFAULT_MODEL_SERVICE_TIERS | Per-model default service tier (default, flex, priority, reserved) | You want a cheaper or a faster tier without changing every client | cost |
MODEL_ALIASES | Maps a name your clients already send onto a model you serve | Clients ask for a model name the gateway does not have | |
COST_TRACKING | Per-request cost estimates in the logs and the Usage API | You want to see spend per request or per user | |
CLOUDWATCH_METRICS | Publishes request, token and cost metrics to CloudWatch | You want dashboards and alarms on gateway traffic | cost |
USAGE_API | Serves the OpenAI-compatible organization usage and cost endpoints | Clients or a billing job read usage back from the gateway | security cost |
LOG_LEVEL | How much the server logs | The default info logs more than you want to ingest, and you only need warnings and above | cost |
OTEL_ENABLED | Emits OpenTelemetry traces (AWS X-Ray and any OTLP backend) | You already run distributed tracing | cost |
CORS_ALLOW_ORIGINS | Browser origins allowed to call the API | A browser application calls the gateway directly | security |
ENABLE_PROXY_HEADERS | Trusts X-Forwarded-* for the client address and scheme | The gateway sits behind an ALB, CloudFront or another proxy | security |
TRUSTED_HOSTS | Host header allowlist | You cannot validate the host at the load balancer | security |
ENABLE_DOCS | Serves the interactive API documentation | You want Swagger UI on a non-public deployment | security |
Tier 3 — Everything else¶
Grouped by page. The defaults work; open a page when a Tier 1 or Tier 2 row above sends you there, or when the reason in the last column is yours.
| Page | Variables | Typical reason to open it |
|---|---|---|
| Regions & AWS clients | 24 | Region pinning, retry and connection tuning, failover backoff, Mantle, data residency |
| Storage | 22 | A separate bucket or prefix, lifecycle rules, vector stores, the shared DynamoDB table |
| Models & routing | 26 | A model name your clients already use, a Marketplace or SageMaker model, per-model defaults |
| Authentication & tenants | 24 | Cognito instead of an API key, per-tenant keys, the discovery documents agents read |
| HTTP server & MCP | 28 | Mounting the routes elsewhere, a browser client, a proxy in front, TLS, MCP tool selection |
| Bedrock features | 28 | Guardrails, stored sessions, a cheaper or faster service tier, the Realtime API, ARN access |
| Observability & usage | 21 | Sending traces or metrics somewhere, per-request cost, the Usage API |
All variables A-Z
Every documented environment variable, with the page that documents it.
Configuration Order¶
When deploying stdapi.ai, configure settings in this recommended order:
- IAM Permissions - Set up AWS access first
- AWS Services and Regions - Choose the Bedrock regions that serve your models, and how requests fail over between them
- Storage - Point file, image, video, batch and vector store features at their S3 buckets
- Authentication - Secure your API with authentication
- Optional features - Add observability, the Bedrock features such as guardrails, and model routing as needed
IAM Permissions¶
Every AWS permission the gateway needs is on the IAM Permissions page: the Amazon Bedrock statements every deployment requires, then one section per optional feature — S3 file storage, vector stores, the shared table, speech, translation, cost tracking, the Usage API, tenant keys — each listing the exact actions and resources it adds. It ends with copy-ready complete policy examples and the AWS tag policy requirements. Go there to write the role your deployment runs under, or when a feature returns an access-denied error.
Deprecated Settings¶
Deprecated and Ignored
TOKENS_ESTIMATION (default: false) and TOKENS_ESTIMATION_DEFAULT_ENCODING (default: None) are deprecated and ignored: tiktoken-based token estimation has been removed from the project. Token counts are now sourced directly from AWS billing data when available. Remove these variables from existing configurations.
Next Steps¶
- IAM Permissions — Complete IAM policy reference
- Authentication & Security — Secure your deployment
- Resilience & Failover — Region routing and failover behavior
- Logging & Monitoring — Observability and metrics