Skip to content

Authentication & Security

stdapi.ai provides flexible authentication options and built-in security mechanisms to protect your API. Choose from API key authentication for simple deployments to identity management delegated to Amazon Cognito, AWS IAM Identity Center or any OIDC provider — all backed by built-in SSRF protection, host header validation, and configurable encryption in transit.

You want to… Read What it costs you
Put one key in front of the whole deployment API Key Authentication A secret to store and rotate; the key is read once at startup, so rotation needs a task replacement
Give each customer or team its own key, scoped and revocable Tenant API Keys A shared DynamoDB table, and one SSM parameter per tenant to collect and delete. Scopes bound invocation, not stored data
Bill a customer's model usage to their AWS account Tenant AWS credentials A cross-account role per tenant; five operation families are refused, and Bedrock Guardrails become unavailable
Stop rotating keys, and identify the caller per request Amazon Cognito User Pool Tokens A user pool and its app clients to run; revoking a caller waits for their current token to expire
Let an AI agent obtain its own credential Authentication Discovery for Agents Two settings, or one with a user pool. The agent still needs an app client created for it in advance
Hand sign-in and SSO to AWS, before the request arrives OIDC, Cognito & IAM Identity Center ALB or API Gateway configuration the Terraform module does not write; stdapi.ai then sees no caller identity
Sign requests with IAM, service to service AWS IAM An API Gateway in front of the deployment; SigV4 on every client
Run with no credential at all No Authentication Network controls become the only defense — a reachable endpoint is an open one, billed to you
Know what protects a request whatever the credential Application Security Nothing: SSRF blocking, Host validation and the size limits are on by default or one setting away
Pass a Security Hub baseline, and monitor threats Security Hub, GuardDuty & DNS Firewall Opt-in module variables, and a dedicated VPC — they have no effect on your own subnet_ids
Encrypt past the load balancer, or demand a client certificate Encryption in Transit Certificates to mount and renew; the Terraform module does not configure end-to-end TLS

Authentication Methods

Method Best for AWS infrastructure Enforced by
API Key Server-to-server, simple deployments Any stdapi.ai
Tenant API keys Serving several customers or teams from one deployment, each key scoped and revocable DynamoDB table stdapi.ai
Amazon Cognito user pool token Per-user access, autonomous agents, revocable credentials Cognito user pool stdapi.ai
OIDC / Cognito / IAM Identity Center User-facing apps, SSO, workforce identity ALB or API Gateway AWS (before request reaches stdapi.ai)
AWS IAM Service-to-service within AWS API Gateway AWS (SigV4)
No Authentication Local development, trusted VPC Any Network controls only

The API key, tenant API keys and Cognito tokens can be used together: AUTHENTICATION_MODE selects which of them a deployment accepts, and defaults to accepting every method that is configured.

When stdapi.ai enforces the credential itself — the API key, tenant API keys and Cognito user pool tokens — a rejected request is answered with 401 Unauthorized, a WWW-Authenticate: Bearer challenge, and a body that states nothing beyond Unauthorized — the reason is written to the server log only, so the response cannot be used to work out which half of a credential was wrong. That challenge is also where an AI agent starts: see Authentication Discovery for Agents. The edge-enforced methods answer an unauthenticated request themselves, before it reaches stdapi.ai, with whatever their own configuration says — commonly a redirect to the identity provider.

API Key Authentication

The API key authentication method uses a single API key that mimics the behavior of upstream OpenAI and Anthropic authentication. It must be provided by clients in the Authorization: Bearer <key> or X-API-Key header.

Three sources are supported — configure exactly one. If more than one is set, the first match in this precedence order is used (the others are ignored): Direct value → SSM Parameter Store → Secrets Manager.

  • Direct value (API_KEY) — for local development and testing only; not recommended for production.
  • SSM Parameter Store (API_KEY_SSM_PARAMETER) — the parameter must already exist before startup.
  • Secrets Manager (API_KEY_SECRETSMANAGER_SECRET) — the secret must already exist; supports a configurable key within the JSON secret via API_KEY_SECRETSMANAGER_KEY.

For the full list of environment variables and required IAM permissions for each method, see the Configuration Guide.

In-Memory Key Protection

stdapi.ai never stores API keys in plain text. At startup, the key is retrieved, salted, and hashed using an industry-standard cryptographic function; only the hash is retained in memory — even a full memory dump cannot reconstruct the original key. Verification uses constant-time comparison to prevent timing-based side-channel attacks.

Terraform Module

The Terraform module offers additional options for providing the API key:

  • api_key — provide the key directly; it is injected as an ECS Secret (sourced from an encrypted SSM parameter with KMS), never exposed in task definitions or logs.
  • api_key_create — auto-generate a secure 64-character random key; the generated value is returned as a sensitive Terraform output.
  • api_key_ssm_parameter / api_key_secretsmanager_secret — reference an existing SSM parameter or Secrets Manager secret; the module does not create these resources.

Amazon Cognito User Pool Tokens

stdapi.ai can accept the bearer tokens issued by an Amazon Cognito user pool as an alternative to the static API key, validating each one on every request. Each caller — a person or an application — gets its own short-lived credential, so there is no shared key to rotate, and the verified caller becomes the identity per-user cost attribution bills against.

Clients send the token exactly as they would send an API key, in either header:

curl https://your-gateway.example.com/v1/models \
  -H "Authorization: Bearer eyJraWQiOi..."

# Anthropic-compatible routes, whose SDK uses its own header
curl https://your-gateway.example.com/anthropic/v1/models \
  -H "x-api-key: eyJraWQiOi..." -H "anthropic-version: 2023-06-01"

What you configure in AWS

  1. A user pool, in any Region. Its ID (for example eu-west-3_a1b2c3d4e) goes into AWS_COGNITO_USER_POOL_ID; the Region is read from the ID itself, and so is the issuer URL — nothing else names it, here or in agent discovery.
  2. One or more app clients in that pool. Their IDs go into AWS_COGNITO_CLIENT_IDS — a token issued to any other app client is rejected, so this list is required.
  3. A pool domain, which is what gives the pool its OAuth 2.0 authorization and token endpoints. It is needed in two cases: to demand a scope such as stdapi/invoke through AWS_COGNITO_REQUIRED_SCOPES, since custom scopes exist only on tokens those endpoints issue and on a resource server that defines them; and for the agent discovery flow, scope or no scope, since an agent has nowhere to obtain a token otherwise. Machine-to-machine clients use the client_credentials grant there.

No IAM permission is involved: the pool's signing keys are public. The task must be able to reach them, though — it reads them over HTTPS from cognito-idp.<region>.amazonaws.com in the pool's Region at startup, and fails to start if it cannot. Give the task outbound HTTPS to that host: NAT or internet egress, or an interface VPC endpoint for com.amazonaws.<region>.cognito-idp in that Region — note that AWS declares a pool with a domain assigned incompatible with that endpoint, so a pool used for scopes or for agent discovery needs the egress path.

The pool decides who can call the gateway

Any identity that can obtain a token from an app client listed in AWS_COGNITO_CLIENT_IDS can call the API, and its model usage is billed to your account. A Cognito user pool allows self sign-up by default, so pointing the gateway at a customer-facing pool without further restriction turns it into an open, self-service one.

Before enabling it: disable self-registration on the pool (AllowAdminCreateUserOnly), or dedicate an app client to the gateway and require one of its resource-server scopes.

What stdapi.ai validates

Every request is checked against all of the following, and any failure returns the same 401 Unauthorized with no detail (the reason is recorded in the server log only):

Check Requirement
Signature RS256, against the pool's published keys. Unsigned (alg=none) and symmetric (HS*) tokens are always rejected.
Issuer Exactly the configured pool's issuer — see AWS_COGNITO_ISSUER_TYPE. A token from another pool is rejected.
Token use An access token. Identity tokens are rejected unless AWS_COGNITO_ACCEPT_ID_TOKEN is enabled.
Application The token's app client is in AWS_COGNITO_CLIENT_IDS.
Validity period Not expired, and not used before its start time, with a one-minute tolerance for clock drift.
Scopes All of AWS_COGNITO_REQUIRED_SCOPES are present.

Signing keys are loaded once, at startup

The pool's public keys are read at startup and kept in memory, so validation adds no network call and no measurable latency to a request. When a pool rotates its keys, the first request carrying a token signed by the new key reloads them, at most once every five minutes — a forged key identifier cannot turn requests into outbound traffic. A deployment that cannot read the keys at startup fails to start rather than serving requests it could not authenticate.

Revocation takes effect when the token expires

A token is validated against the pool's published keys and its own claims, without calling the pool, so nothing is checked back with it once it has been issued. Disabling a user, deleting them or signing them out therefore stops the next token, not the one already in their hands: that one keeps working until its exp, up to the app client's access-token validity (one hour by default, 24 hours at most). Keep that validity short — it is the upper bound on how long a revoked credential stays usable.

Username and password sign-in yields no custom scope

Tokens obtained by signing in directly against the user pool API carry the single scope aws.cognito.signin.user.admin. Custom scopes only exist on tokens issued by the pool's OAuth 2.0 token endpoint. Requiring stdapi/invoke therefore rejects every client that signs in with a username and password — leave AWS_COGNITO_REQUIRED_SCOPES empty unless all your clients use the OAuth 2.0 endpoints.

Both methods, or one

With both a user pool and an API key configured, either credential is accepted: a bearer value shaped like a signed token is validated against the pool, anything else is compared to the API key. Set AUTHENTICATION_MODE to cognito or api_key to accept only one of them — the deployment then refuses to start if the other one is configured too, so a credential is never accepted by accident.

Tenant API Keys

One deployment can serve several customers or teams, each holding an API key of its own — shaped sk-std-<key id>-<secret> — that is validated on every request and scoped to the models and endpoints its tenant is entitled to. Enable the method with TENANT_API_KEYS; it needs the shared DynamoDB table. Keys are delivered under a delivery prefix that has a default and only needs overriding on a deployment that shares an AWS account with another. Clients send the key like any API key, in the Authorization: Bearer <key> or X-API-Key header. The deployment-wide API key and Cognito tokens keep working unchanged alongside it — enabling tenant keys changes nothing for existing credentials.

Declaring tenants and receiving their keys

A tenant is a record in the shared table: its name, its scopes, and a disabled flag. The Terraform module declares one per entry of its tenants variable:

tenants = {
  "acme" = {
    models_allow    = ["anthropic.*", "amazon.nova-lite-v1:0"]
    endpoints_allow = ["/v1/chat/completions", "/v1/models"]
  }
  "globex" = {
    models_deny = ["*opus*"]
  }
}

Scopes bound the models a tenant may invoke and the endpoints it may call; they do not partition stored objects, which stay deployment-wide (the caveat below).

The key secret never enters Terraform state. Terraform owns the tenant record only; the server notices a declared tenant that has no credential yet — at startup and once a minute — mints a 256-bit secret for it, stores only a salted hash in the table, and delivers the full key exactly once as an SSM SecureString parameter named <prefix>/<key id> (the module's tenant_keys output gives the exact name per tenant). Retrieve it, hand it to the tenant, then delete the parameter:

aws ssm get-parameter --name /my-deployment/tenant-keys/AbC123... \
  --with-decryption --query Parameter.Value --output text

Where keys must be rotated, the server stores each key in an AWS Secrets Manager secret of its own instead — see Rotating tenant keys below.

Only a hash is ever stored

The table holds a salted BLAKE2b-256 digest of the secret, compared in constant time on every request — the same in-memory protection the deployment API key gets. Neither the table, nor Terraform state, nor a backup of either can reconstruct a tenant's key; the only copy is the delivered parameter, which is yours to delete after delivery — or, with the Secrets Manager store, the secret the tenant reads.

The delivery prefix is a trust boundary in both directions

The parameter is created once and never overwritten, which is what makes minting idempotent across instances — so whoever creates it defines the secret. A principal able to call ssm:PutParameter under the prefix can therefore pre-create <prefix>/<key id> for a tenant that does not exist yet and hold a valid key from the moment it is declared, exactly as read access there exposes the keys already delivered. Grant both actions on <prefix>/* to the deployment's task role and to the operators who collect the keys, and to nothing else.

Read access is wider than that grant while TENANT_KEY_SSM_KMS_KEY_ID is unset: the parameter is then encrypted with the AWS-managed alias/aws/ssm key, whose key policy lets any principal of the account decrypt through Parameter Store, so ssm:GetParameter on the path is the only permission an intruder needs. Point the setting at a key of your own and reading a delivered key also requires kms:Decrypt on it — the Terraform module points it at the deployment's KMS key automatically.

Without the Terraform module

Any tool that can write a DynamoDB item can declare a tenant: pk = TENANT, sk = tenant#<key id> (a key ID is 16 letters and digits of your choice, unique per tenant), attributes name (string), schema (number, 1), optional disabled (boolean), the four scope lists below as lists of strings, and the optional rate limits requests_per_minute and tokens_per_minute as numbers — whole, at least 1; a string or a zero refuses the key rather than reading as unlimited. The server mints and delivers the key the same way.

Scopes

Each tenant record may carry four pattern lists, matched with * and ? globs. A list that is absent restricts nothing; a list that is present and empty allows nothing; a deny match always wins over an allow match.

List Matched against A refused request answers
models_allow / models_deny The resolved model ID, after aliases, wildcards and deprecation fallbacks — an alias cannot launder a denied model The standard 404 model_not_found, indistinguishable from a model that does not exist
endpoints_allow / endpoints_deny The matched route's path template, e.g. /v1/chat/completions or /v1/files/{file_id} The same detail-free 401 Unauthorized as any refused credential

Deny lists match names, so prefer an allow list

Model patterns are matched against the model ID the request resolves to. With AWS_BEDROCK_ALLOW_MARKETPLACE_ENDPOINT_ARN enabled, an endpoint addressed by its ARN keeps that ARN as its ID, which a name-based pattern such as mistral.* does not match. models_allow fails closed on it — no pattern matches, so the model is refused — while a deny-only tenant would reach it: add arn:* to models_deny where that opt-in is on.

A release that adds endpoints widens an endpoint deny list

endpoints_deny is matched against the path templates the running version serves. Releases add routes — a new API dialect mounts a whole family of them at once, as /api/* did for the Ollama API — and a deny list written against the previous version matches none of them, so a tenant scoped by denial silently gains them on upgrade. endpoints_allow fails closed on the same upgrade: a path it does not list stays refused. Prefer an allow list here too, and re-read the API reference after an upgrade whenever you keep a deny list.

Both are enforced at choke points every request passes through — the authentication dependency for endpoints, the single model-resolution step for models — not per route, so a new endpoint or model route cannot bypass them. GET /v1/models is deliberately not filtered per tenant: the catalogue advertises what the deployment serves, and the invocation-time check is the authority. A Realtime client secret minted with a tenant key stays bound to that tenant: the session it opens carries the same scopes and stops resuming once the key is revoked or disabled.

Scopes bound what a tenant may invoke, not what it may read

Stored objects carry no tenant ownership. Files, vector stores, batches and their result files, conversations, stored responses, and the Anthropic files and message batches are all deployment-wide: neither list filters them, so a tenant allowed on /v1/files lists, downloads and deletes what every other tenant uploaded. The model and endpoint scopes bound what a key may invoke and call; they are not a data boundary.

Where tenants must not reach one another's data, deny the storing endpoints and keep the deployment stateless for them:

endpoints_deny = [
  "*/v1/files*",
  "*/v1/vector_stores*",
  "*/v1/batches*",
  "*/v1/messages/batches*",
  "*/v1/conversations*",
  "*/v1/responses/*",
]

The leading * covers the routes prefix each dialect is mounted under, such as ANTHROPIC_ROUTES_PREFIX. /v1/responses itself stays allowed so the Responses API keeps serving requests; a response the tenant asks to store is still written, only no longer readable by it. Mutually untrusted tenants that need any of these features want one deployment each — a separate bucket and table is the only isolation there is.

Validation, caching and revocation

Validation is a direct read of the tenant's two records, cached in each server instance for TENANT_KEY_CACHE_SECONDS — 60 seconds by default. That cache is the revocation window: a key that is revoked, disabled or re-scoped keeps its previous decision for up to 60 seconds per instance, and no longer. Revoke a key by removing its tenant from tenants (destroying the record), or suspend it by setting disabled = true — both also end the overlap during which a rotated key's predecessor still works. Unknown key IDs are negative-cached, bounded in size and time, so a flood of fabricated keys neither amplifies table reads nor grows memory; a key that does not match its cached entry is re-read at most once a second, so a freshly rotated key is accepted everywhere within one read.

The table being unreachable fails closed

When tenant keys are enabled but the DynamoDB table cannot be read, a tenant-shaped credential is refused with 503 — never accepted, and never turned into a 401 that would mislabel a valid key as wrong. The reason (the IAM action, the table) is written to the server log. Other credential kinds are unaffected.

A tenant key and a user token together

A request may carry both X-API-Key: sk-std-... (the tenant key) and Authorization: Bearer <token> (a Cognito user token): both are then verified — the tenant key authorizes and scopes the request, the token identifies the user for per-user cost attribution. This is the one place tenant keys extend the header rules: for every other combination, X-API-Key keeps winning exactly as before. Without a user, the tenant's key ID is the identity the request is attributed to.

Rate limits per tenant

Each tenant key may be capped in requests per minute and tokens per minute, on its record or for every tenant at once with TENANT_RATE_LIMIT_REQUESTS_PER_MINUTE and TENANT_RATE_LIMIT_TOKENS_PER_MINUTE; a record's own value wins over the deployment default. Nothing is counted, and no table write happens, until a limit is declared somewhere.

tenant_rate_limit_requests_per_minute = 600      # every tenant, unless it says otherwise

tenants = {
  "acme" = {
    requests_per_minute = 6000
    tokens_per_minute   = 2000000
  }
}

The limits are enforced across every instance of the deployment: each minute is a fixed window aligned to the clock, counted once in the shared table, so the instances never admit more than the request limit between them — and, as the reservation below explains, possibly less. The counter is kept off the steady-state request path, but not off all of it: an instance reserves request slots in batches that start at one and double within the minute, capped at an eighth of the request limit, and a request finding no slot left waits for one table write before it is admitted. The first two requests of each minute for a key on an instance therefore wait, as does a later one whenever the background refill has not landed yet — and below a request limit of 16 the cap is a single slot, so every request of such a tenant waits for a write. A tenant limited on tokens alone waits once a minute, for the request that reads the window's totals. A slow or throttled table shows as latency on exactly those requests. On HTTP routes a key over its limit answers the vendors' own refusal — 429 with error type rate_limit_error (code rate_limit_exceeded on OpenAI-compatible routes) and a retry-after naming the seconds left in the minute, which the OpenAI and Anthropic SDKs honour automatically. Every response of a limited tenant on an OpenAI- or Anthropic-compatible route also carries its dialect's rate-limit headers, one triple per declared limit: x-ratelimit-limit-requests, x-ratelimit-remaining-requests and x-ratelimit-reset-requests with their -tokens counterparts on the former, anthropic-ratelimit-requests-* and anthropic-ratelimit-tokens-* on the latter. A tenant capped on requests alone publishes the requests triple only, one capped on tokens alone the tokens triple only; the six appear together only where both limits are declared. The Cohere- and Ollama-compatible routes are limited identically but publish no rate-limit headers and no error type: their refusal is that dialect's own plain error envelope with status 429 and the retry-after header alone.

Tokens count what the model billed — input, cache-write and output tokens; cached reads are free, as on the Anthropic API. Because that number is only known once the model has answered, a request is admitted on an estimate and the actual figure is reconciled afterwards, including the trailing usage of a streamed response and each turn of a Realtime session. The estimate is learned per instance from the key's own traffic. A key an instance has not billed yet is charged an eighth of its token limit per request in flight: a cold instance — every one, after a deployment or a scale-out — admits about eight concurrent requests of that key whatever their real size, and refuses the ninth with a 429 naming the token limit while the minute's total still reads zero; raising the limit does not change that ratio. As soon as one request bills there, each request in flight is charged the key's mean tokens per billed request on that instance instead, capped at the limit — a Realtime session counts as one request in every minute it bills in. A minute can therefore exceed its token limit: a request in flight is charged the mean rather than what it will bill, so a burst of requests larger than the mean overshoots by the difference, and billed tokens reach the shared counter about once a second per instance, so a minute can overshoot by a second of every instance's traffic on top. Because the mean is taught by the key's own traffic, a token limit alone also caps the requests a key holds in flight — admitted and not yet billed — at 64 per instance, refused with the same 429; declaring a request limit replaces that ceiling with the request limit. The overshoot is bounded by the requests in flight times the largest per-request token count, so pair a token limit with a request limit sized for the concurrency the tenant needs, and size the pair for the largest per-request token count. The next requests are refused until the minute ends, and the overshoot is forgiven at the boundary.

What is and is not limited

  • Only tenant keys. The deployment API key and Cognito tokens carry no per-key budget; rate them at the load balancer — the Terraform module's alb_waf_rate_limit does so per client address.
  • A Realtime session counts once, at the handshake, and is never cut mid-session: the turns it bills are reconciled against the token limit and refuse the tenant's next connection, not the open one. A session opened with an ephemeral client secret costs a second request, for the POST /v1/realtime/client_secrets that mints it.
  • A refused Realtime handshake is not a 429: the WebSocket upgrade succeeds (101) before the limit is checked, then the server sends one terminal error event — type invalid_request_error, code rate_limit_exceeded, the message naming the seconds to wait — and closes the socket with code 3000, reason invalid_request_error.rate_limit_exceeded. No retry-after and no rate-limit header reach a WebSocket client, so a Realtime client must back off on its own; the 429 appears only in the server request log.
  • Batch jobs are not counted: their output is produced outside any request.
  • The minute is a fixed window, so a client can spend one budget at the end of a minute and the next at its start — up to twice the limit within sixty seconds across the boundary.
  • A tenant spread over many instances can be refused below its request limit. Each instance reserves request slots ahead in batches of up to an eighth of the limit, charged to the shared counter at once, and never returns what it does not use before the minute ends. An instance that stops seeing the tenant's traffic mid-minute — a keep-alive connection pinned elsewhere, an idle client, a scale event or a rolling deployment — strands up to about a fifth of the limit; a request the token limit refuses consumes no slot, which stays with the instance that reserved it. The shortfall grows with the instance count and with bursty or short-lived traffic: declare the limit with headroom for the instances it is spread over, or keep a tenant's traffic on fewer of them. The same reservation makes the remaining headers one instance's view, so two instances may report different figures for the same key.

The counter being unavailable fails closed

When a limited tenant's counter cannot be written — the dynamodb:UpdateItem permission missing, the table throttled or unreachable — the instance keeps serving the requests it had already reserved, then refuses the tenant's next ones with 503 (529 on Anthropic-compatible routes; on a Realtime WebSocket, an error event of type server_error after the upgrade) rather than a 429 that would mislabel them as over the limit, and never serves them unlimited. A table that stalls instead of answering is given up on within seconds — the wait is a short timeout and few retries, not the minutes a model is allowed — so the refusal is prompt, for the request waiting on the write and for every request of that key queued behind it. The reserved grace is a request limit's: a tenant limited on tokens alone holds no reservation and is refused as soon as the counter cannot be written. The action, the table and the region are in the server log, and a deployment declaring any limit checks the permission at startup. Tenants without a limit are unaffected.

Rotating tenant keys

A key delivered through Parameter Store is delivered once and never changes: rotating it means destroying the tenant and declaring it again, under a new key ID. Set TENANT_KEY_SECRETSMANAGER_PREFIX and the server stores each key durably instead, as the current version (AWSCURRENT) of an AWS Secrets Manager secret named <prefix>/<key id>, and rotates it in place. The Terraform module switches to this store as soon as tenant_key_rotation_days is set or any tenant declares key_generation, creates one secret per tenant on the deployment's KMS key, and names it in the tenant_keys output.

A rotation happens without an operator in the loop, on two triggers that combine:

  • On a schedule — TENANT_KEY_ROTATION_DAYS: every key older than that, counted from its mint or its last rotation, is rotated by the next reconciliation pass (at startup, then once a minute).
  • On demand — raise key_generation on the tenant record (tenants["acme"].key_generation = 2 in the module, or the attribute written directly): the key is rotated once, when the value exceeds the generation recorded with the secret, and never again for the same value. Declarative and idempotent, so a plan that bumps it is safe to apply twice.
tenant_key_rotation_days = 90

tenants = {
  "acme" = {
    models_allow   = ["anthropic.*"]
    key_generation = 2   # raise to rotate now
  }
}

The new key is written as the secret's pending version, accepted by the gateway, then promoted to AWSCURRENT; the superseded key becomes AWSPREVIOUS and keeps authenticating for TENANT_KEY_ROTATION_OVERLAP_SECONDS — 7 days by default — so a client that re-reads its secret on its own cadence is never locked out. Only the last superseded key is kept, so rotating again inside that window retires the older one at once. A client that re-reads AWSCURRENT right after the rotation is accepted on every instance within one read. Set the overlap to 0 for a hard cutover; whatever the overlap, setting disabled = true on the tenant refuses both keys within TENANT_KEY_CACHE_SECONDS.

A tenant re-reads its own key with a policy granting it secretsmanager:GetSecretValue on its own secret alone (and kms:Decrypt on the deployment's key, through Secrets Manager); the secret's ARN is in the module's tenant_keys output. That grant is yours to write — the module does not attach a resource policy to the secret.

The secret prefix is a trust boundary in both directions

Exactly as with the delivery prefix: whoever writes a version defines the key. A principal able to call secretsmanager:PutSecretValue under the prefix can stage a well-formed key as a tenant's next version and hold it from the moment the rotation adopts it, and read access there exposes every tenant's current and previous key. Grant the write actions to the deployment's task role and to nothing else; grant each tenant read access to its own secret only.

What the server owns, and what it does not

The server creates a tenant's secret when none exists and writes every version, so a deployment without the module works the same. It never deletes a secret and never tags one: the module creates each secret with the deployment's tags and destroys it with the tenant record, and a deployment declaring tenants by hand deletes the secret when it destroys the tenant — the revocation names it in the server log. A key minted through Parameter Store before the store was enabled keeps working unchanged and receives a secret of its own at its first rotation.

AWS Security Hub

The rotation is driven by the gateway, not by a rotation function, so the secrets carry no rotation configuration and the SecretsManager.1 control ("secrets should have automatic rotation enabled") reports them as failed — a control that only a Lambda-based rotation can pass. The periodic-rotation control (SecretsManager.4) passes while TENANT_KEY_ROTATION_DAYS is 90 or less, and the unused-secret control (SecretsManager.3) flags a secret no tenant has read for 90 days, which is expected for a tenant that keeps its key locally.

A mixed fleet during a rolling deployment

Instances still on a release without rotation keep validating a rotated tenant's current key — its hash is stored where they read it — but do not honour the overlap: on those instances the superseded key is refused as soon as the new one is recorded. Instances on a release with rotation but older than v1.18.0 start the overlap at that same moment rather than at the promotion, so a promotion that has not landed locks the tenant out on those instances alone. Roll the fleet before the first rotation is due, or disable the schedule until it is rolled.

Tenant AWS credentials — the tenant's own quota and bill

A tenant may go one step further and bring its own AWS account: register an IAM role of that account against its key, and every model invocation made with the key runs under that role — the tenant's own Amazon Bedrock model access, quotas, throttling and bill, instead of the deployment's. Enable it with TENANT_AWS_CREDENTIALS and declare the role on the tenant record:

tenants = {
  "acme" = {
    aws_role_arn = "arn:aws:iam::210987654321:role/stdapi-acme"
  }
}

The design is AWS's own cross-account confused-deputy pattern: the server assumes the tenant's role with sts:AssumeRole, presenting an ExternalId the server mints for that tenant — read it from the external_id attribute of the tenant's secret#<key id> record, the only place it is published. The tenant writes it into the role's trust policy:

{
  "Effect": "Allow",
  "Principal": { "AWS": "arn:aws:iam::<deployment account>:root" },
  "Action": "sts:AssumeRole",
  "Condition": { "StringEquals": { "sts:ExternalId": "<minted external id>" } }
}

The role's permission policy is the tenant's to scope — bedrock:InvokeModel* and the Converse actions on the models it wants to allow. No secret exists anywhere at rest: the stored role ARN is public-by-design and the ExternalId is only meaningful when presented by the deployment's own IAM principal. The tenant revokes the grant at any moment by editing its own trust policy; sessions are capped at one hour by AWS role chaining, so a revocation is fully effective within that hour, and the server never signs a new request with a session it could not reopen.

What runs on whose account. The tenant's credential covers model invocations only — the Converse and InvokeModel families, streaming included. That includes the embedding calls behind vector store indexing and search: the tenant's role must allow the deployment's configured VECTOR_STORE_EMBEDDING_MODEL for vector store operations to work, and their embedding spend is the tenant's. Everything else a request may touch stays on the deployment's account: Amazon Polly, Transcribe, Translate and Comprehend, S3 storage and Knowledge Bases. Five operations are refused for a credential-carrying key rather than silently billed to the deployment, each with a clear error naming the reason: the Batch API (a job outlives the one-hour session), real-time speech-to-speech model sessions (a bidirectional stream signs once, as the server), asynchronous (video) generation (the job writes to the deployment's bucket), reranking models (their per-query invocations run through a service the tenant's credential cannot sign), and models only served by Amazon Bedrock Mantle, by a Bedrock Marketplace endpoint or by a SageMaker AI endpoint. A model served by Mantle by default but also available on the classic runtime — the GPT-5.6 and GPT-6 families — is transparently served from the runtime instead, where the tenant's credential signs and pays. Vector store indexing for such a key runs in the server that accepted the file even when the deployment hands indexing to a queue, so it stays on the tenant's bill.

Not compatible with Amazon Bedrock Guardrails

A guardrail configured in the deployment's account cannot be evaluated by a tenant's principal — AWS Guardrails have no cross-account path. Rather than silently serving tenant requests unguarded, the server refuses to start when TENANT_AWS_CREDENTIALS is enabled while AWS_BEDROCK_GUARDRAIL_IDENTIFIER or a model-alias guardrail is configured. A tenant may still name a guardrail of its own account per request where header overrides are allowed — on the chat APIs, where the guardrail rides the model invocation itself, it is evaluated by the tenant's principal, in the tenant's account. Every other route applies a guardrail as a separate evaluation under the deployment's identity, which cannot read a guardrail of the tenant's account: the request fails there rather than serving unguarded content.

Why a role, and nothing else

Cross-account roles are the only credential form served, by policy and not by omission. Amazon Bedrock API keys are refused: short-term keys last at most 12 hours — unstorable by construction — and long-term keys are IAM-user bearer tokens AWS itself marks "for exploration only"; storing them would turn the tenant table into a vault of live credentials. Raw access keys are long-lived secrets at rest with no tenant-side revocation story. OIDC federation adds an identity provider to operate without removing any of the above. A role plus ExternalId stores nothing worth stealing.

A few behaviors worth knowing:

  • The model catalogue is not per-tenant. GET /v1/models advertises what the deployment serves, from the deployment's account. A model the tenant's account has no access to fails at invocation with an honest 403: "Your AWS account does not have access to this model."
  • Errors never name the deployment's account. A revoked trust policy, a wrong ExternalId or a deleted role answers a fixed 403 — "The AWS credential registered for this API key could not be used…" — with the full AWS detail written to the server log only. AWS STS being throttled or unreachable is not the tenant's registration and answers a retryable 503 instead, so a client retries it rather than auditing a trust policy that is fine.
  • A tenant's throttle is its own. Tenant requests still fail over across the configured Bedrock Regions, but a throttle of the tenant's account is answered as the tenant's own 429 and never marks a Region as blocked for other callers — Bedrock quota is per account, per Region. One exception, and it is opt-in: with AWS_ADAPTIVE_RETRY enabled, the retry token bucket belongs to the shared client rather than to a credential, so a tenant being throttled heavily can slow other callers' retries. Leave adaptive retry off — the default — where tenants must not affect one another.
  • Cross-Region inference profiles route within the tenant's account: destination Regions that are opt-in must be enabled in the tenant's account for the profile to serve them.
  • Cost reporting stays honest. Usage billed to the tenant's account is marked billed_to: tenant in the usage log and is never priced into the deployment's cost totals.

Authentication Discovery for Agents

An AI agent that meets an API it has never been configured for has one thing to go on: the 401 Unauthorized it just received. stdapi.ai answers that request with everything the agent needs to obtain a credential on its own — the standard discovery flow behind MCP client authentication.

Name the public URL and the token issuer, and the discovery surface turns on:

Setting Value
OAUTH_RESOURCE_IDENTIFIER The public URL clients dial, exactly as they dial it — for example https://api.example.com
OAUTH_AUTHORIZATION_SERVERS The issuer URL of whatever issues your tokens. Leave it unset with a user pool configured: its issuer is published for you
OAUTH_SCOPES_SUPPORTED Leave it unset too — AWS_COGNITO_REQUIRED_SCOPES is published, since an agent that is not told which scope to ask for obtains a token without it and is refused on every retry

With a user pool, one variable turns discovery on

The pool named by AWS_COGNITO_USER_POOL_ID issues the tokens this deployment accepts, so it is the authorization server clients must be sent to, and the scopes it requires are the scopes they must ask for. Both are published from it, and the issuer is exactly what the pool signs into iss:

https://cognito-idp.<region>.amazonaws.com/<pool-id>
https://issuer-cognito-idp.<region>.amazonaws.com/<pool-id>   # AWS_COGNITO_ISSUER_TYPE=updated

<region> and <pool-id> are read from the pool ID itself, and the host follows that Region's AWS partition — amazonaws.com.cn in China, amazonaws.eu in the European Sovereign Cloud. Set the two variables only for an issuer no pool can supply, such as an identity-aware proxy in front of the deployment; the pool's own issuer must stay in the list, or startup fails rather than sending clients to an authorization server whose tokens are refused.

What an agent then sees:

$ curl -i https://api.example.com/v1/models
HTTP/1.1 401 Unauthorized
WWW-Authenticate: Bearer resource_metadata="https://api.example.com/.well-known/oauth-protected-resource", scope="stdapi/invoke"

$ curl https://api.example.com/.well-known/oauth-protected-resource
{
  "resource": "https://api.example.com",
  "authorization_servers": ["https://cognito-idp.eu-west-3.amazonaws.com/eu-west-3_a1b2c3d4e"],
  "scopes_supported": ["stdapi/invoke"],
  "bearer_methods_supported": ["header"],
  "resource_name": "stdapi.ai (Community Edition)",
  "resource_documentation": "https://stdapi.ai/api_reference/"
}

From there the agent reads the authorization server's own metadata, signs in, and retries the request with a bearer token. The document is public and unauthenticated by design — it is read by a client that has no credential yet — and contains nothing beyond the issuer URL you configured. It is also cacheable, and it is advertised from the API catalog and the root Link header, so an agent that starts from either entry point finds it without a 401.

Both the API and the MCP server are covered

The identifier names the deployment's origin, which every surface shares: an MCP client connecting to /mcp or /sse, and an SDK calling /v1/…, all match the same published resource. Only one document is served, at the root.

/.well-known/openid-configuration is deliberately not served

stdapi.ai is a resource server, not an authorization server: it validates tokens, it does not issue them. Mirroring or redirecting the authorization server's own discovery document would create a second, staler copy of something the issuer already publishes authoritatively. Agents reach it through authorization_servers instead — which is exactly what the standard flow does.

Amazon Cognito needs a pre-registered app client

A Cognito user pool publishes neither a dynamic client registration endpoint nor client-id metadata document support, so an agent cannot register itself. Create the app client in the pool and give the agent its client ID (and secret, for confidential clients) ahead of time. Discovery still saves the agent everything else: the issuer, the endpoints, and the scopes.

The pool also needs a domain: its authorization and token endpoints exist only once one is assigned, so without it the issuer's own discovery document names no endpoint the agent can obtain a token from, and the flow dead-ends after the 401.

Browser-hosted clients need CORS

A client running in a browser reads this document cross-origin. Add its origin to CORS_ALLOW_ORIGINS, which it also needs for the API calls that follow.

With a CORS origin configured, the gateway also marks WWW-Authenticate as readable cross-origin, so a page-hosted client reads the metadata URL and the advertised scope straight off the 401. Without one, the browser hides that header and the client falls back to the default well-known path, losing the scope the challenge carries.

OIDC, Cognito & IAM Identity Center

For user-facing applications or enterprise SSO, you can offload authentication to an OpenID Connect (OIDC) provider, Amazon Cognito, or AWS IAM Identity Center (the AWS-native workforce SSO service for centralized employee and partner access).

When authentication is handled at this layer, requests are fully validated before they reach stdapi.ai — the application receives only authenticated, pre-authorized requests. This remains a valid alternative to validating Cognito tokens in stdapi.ai: choose the edge when you need a hosted sign-in flow or session cookies, and choose in-application validation when clients already hold a bearer token and you want the verified caller to drive per-user cost attribution.

AWS-operated identity — no custom implementation to maintain

Cognito and IAM Identity Center are the same identity systems AWS teams already rely on for console access and internal applications. By delegating authentication to these services, you get MFA, SSO, and fine-grained permission sets without building or maintaining custom user management logic — and without the security risk of a home-grown implementation. See Eliminate key rotation with AWS native auth for the operational payoff.

Terraform Module

The Terraform module does not configure the ALB or API Gateway integrations described in this section, nor the identity provider itself — these are set up directly on the ALB or API Gateway. Validating Cognito tokens in stdapi.ai needs no infrastructure beyond the pool: see Amazon Cognito User Pool Tokens.

via Application Load Balancer (ALB)

The AWS ALB can authenticate users before forwarding requests to stdapi.ai. This is the most common way to add OIDC or Cognito authentication.

  • Capabilities: Integrates with any OIDC-compliant identity provider — including AWS IAM Identity Center (workforce SSO), Amazon Cognito User Pools, or third-party providers such as Okta, Auth0, and Google.
  • IAM Identity Center: Expose the Identity Center OIDC application endpoints as the ALB OIDC configuration. This gives employees and partners SSO using their AWS-managed identity, with support for MFA and permission sets defined in your AWS Organization.
  • Documentation: Authenticate users using an Application Load Balancer

via API Gateway

Both REST and HTTP APIs support OIDC and Cognito integration.

AWS IAM

For service-to-service communication within AWS or when using IAM roles for access control, you can use AWS IAM authentication via API Gateway. Clients sign their requests using the AWS Signature Version 4 (SigV4) process, enabling fine-grained access control using IAM policies attached to users or roles.

via API Gateway

Amazon API Gateway provides native IAM authentication for both REST and HTTP APIs.

No Authentication

In certain environments, you may choose to disable application-level authentication and rely entirely on network security.

When to Use

  • Local development or testing using Docker.
  • Internal VPC deployments where traffic is trusted.
  • When an upstream proxy handles all security and identity.

Configuration

To disable application-level authentication, configure none of the four credential sources: no API key environment variable (API_KEY, API_KEY_SSM_PARAMETER, or API_KEY_SECRETSMANAGER_SECRET), no user pool (AWS_COGNITO_USER_POOL_ID), and TENANT_API_KEYS left at false — tenant keys are a credential source on a par with the others, so leaving them enabled keeps every request needing a valid tenant key.

When no authentication method is configured, stdapi.ai will:

  1. Accept all incoming requests without validating an API key.
  2. Log a security warning at startup.

Network Security Requirement

When API key authentication is disabled, you must ensure strict network security. Use Security Groups to restrict access so that the stdapi.ai service only accepts traffic from your trusted Load Balancer or API Gateway.


Application Security

stdapi.ai includes built-in security mechanisms that are active regardless of the authentication method chosen.

Hardened Container Image

The AWS Marketplace image is built on a security-hardened minimal base: no shell, no package manager, no unnecessary system tools, and a distribution rebuilt continuously to keep its known-vulnerability count at or near zero. The community image is built on a standard Debian base — the same application, the same non-root user, a larger base with the system tools a general-purpose distribution ships. Both run as an unprivileged user, and the Terraform module further enforces a read-only root filesystem and drops all Linux capabilities from the ECS task definition, reducing the attack surface to the minimum required for operation.

Subscribe on AWS Marketplace

Supply Chain Security

Supply chain attacks — where malicious code is injected into a software package before it reaches your infrastructure — represent one of the most dangerous threat vectors for self-hosted AI gateways. A compromised package can silently exfiltrate API keys and credentials from every deployment that installs it, and attacks of this kind have targeted the AI tooling ecosystem.

stdapi.ai's distribution model eliminates this attack surface entirely:

  • Container-only distribution — stdapi.ai is never distributed as a pip package or PyPI dependency. There is no pip install stdapi.ai, no transitive dependency chain to compromise, and no package registry to hijack. The only distribution channels are the AWS Marketplace ECR registry (commercial) and GHCR (community image).
  • AWS Marketplace security validation — the commercial container image is scanned and validated by AWS before being made available in the Marketplace. The image you deploy is exactly the image AWS validated — nothing can be inserted between validation and deployment.
  • Minimal base image (commercial) — the Marketplace container has no shell, no package manager, and no unnecessary system tools. There is no mechanism to install additional packages or execute arbitrary commands at runtime. The community image uses a standard Debian base and keeps those tools.
  • Immutable deployment — the container image is pulled once at task start and run as-is. There are no runtime pip install, apt install, or dependency resolution steps that could fetch and execute untrusted code.
  • Read-only root filesystem — when deployed via the Terraform module, the ECS task definition enforces a read-only root filesystem, preventing any modification to the container files at runtime.
  • Temporary credentials with least privilege — CI/CD pipelines that build and publish the container image use short-lived, role-scoped credentials with only the permissions required for that specific job. No long-lived access keys are used in the build or release process. Jobs that do not require AWS access — such as linters and security scanners — run on isolated runners with no AWS credentials at all, limiting the blast radius of any compromised job.

SSRF Protection

Users can pass URLs to the API as multimodal content references (images, documents, audio). The application fetches the content at those URLs before forwarding it to the model. Without protection, a malicious user could supply URLs pointing to internal services — such as the AWS EC2/ECS metadata endpoint or private network resources — and read sensitive data through the model response.

stdapi.ai validates every user-supplied URL via DNS resolution before fetching its content — each hostname is resolved and every resulting IP is checked against the blocklist. This protects against DNS rebinding attacks where a seemingly-safe domain resolves to a blocked address.

  • Baseline Protection: Always-active blocking of loopback (127.0.0.1, localhost), link-local (169.254.169.254), and reserved address ranges. This prevents access to the AWS EC2/ECS metadata service from within the container.
  • Private Network Blocking: By default, the service also blocks every address that is not globally reachable on the public Internet — RFC 1918 networks (10.x.x.x, 172.16-31.x.x, 192.168.x.x), IPv6 unique local addresses, RFC 6598 shared address space (100.64.0.0/10, used by EKS custom networking and Hybrid Nodes) and the remaining special-purpose ranges, in their IPv4-mapped IPv6 form as well. Disable only in controlled environments where accessing local networks is required (SSRF_PROTECTION_BLOCK_PRIVATE_NETWORKS=false).
  • Defense in Depth: Even with SSRF protection enabled, restrict outbound Security Group rules to only the necessary AWS service endpoints.

Network-layer complement: Route 53 Resolver DNS Firewall

SSRF protection blocks by IP range — it doesn't know whether a public IP belongs to a malicious domain (malware, phishing, botnet C2). The Terraform module can close that gap: set dns_firewall_enabled = true to block outbound DNS resolution of known-malicious domains before a fetch ever happens. See AWS Security Hub, GuardDuty & DNS Firewall Integration.

Host Header Validation

To protect against Host header injection and web cache poisoning, stdapi.ai can validate the Host header of incoming requests. When TRUSTED_HOSTS is not configured, no Host header validation is performed.

  • Recommended Approach: Use AWS ALB host-based routing rules to reject invalid Host headers before they reach the application. This is more performant and centrally managed.
  • Application Validation: Use TRUSTED_HOSTS to define a list of approved hostnames (supports wildcards like *.example.com). Requests with non-matching headers are rejected with an HTTP 400 error.
  • Probe Impact: Validation covers /health too. The container image's health probe derives its Host header from TRUSTED_HOSTS and stays green on its own, but a load balancer health check sends the target's IP address as the Host and is rejected — see TRUSTED_HOSTS before enabling it behind an ALB or NLB.

Browser Security (CORS)

If your API is accessed directly from web browsers, Cross-Origin Resource Sharing (CORS) must be configured to prevent unauthorized cross-site requests. When CORS_ALLOW_ORIGINS is not configured, the CORS middleware is not enabled and browser cross-origin requests are blocked by default.

  • Principle of Least Privilege: Only allow specific, trusted domains using CORS_ALLOW_ORIGINS.
  • Offload to Infrastructure: For production, consider managing CORS at the AWS ALB or API Gateway level for better performance and to keep your application logic clean.

Proxy Headers

When running behind a load balancer (ALB) or reverse proxy, stdapi.ai can be configured to trust X-Forwarded-* headers (X-Forwarded-For, X-Forwarded-Proto, X-Forwarded-Port) to correctly identify the original client IP and protocol. When enabled, the real client IP appears in request logs instead of the proxy's IP.

  • Security Warning: Only enable ENABLE_PROXY_HEADERS when the service is deployed behind a trusted proxy. Enabling this in a publicly exposed environment without a proxy allows attackers to spoof their IP address by sending custom headers.

Request Size & Resource Limits

Large or numerous inputs are a denial-of-service vector: a single request can carry hundreds of megabytes of inline base64 content, reference many remote files, or stream for a long time. stdapi.ai provides application-level controls, but the most effective size limit is enforced at the edge, before the request reaches the application.

Application-level controls:

  • Inline & downloaded input size: MAX_INPUT_FILE_SIZE caps the bytes of any single file loaded into memory for model input (base64, data: URIs, and HTTP(S)/S3 sources read for the model). Requests over the limit are rejected with HTTP 413 before the content is fully decoded or downloaded — a spoofed Content-Length cannot bypass it. Disabled by default; set a value aligned with your largest expected input when exposed to untrusted clients.
  • Concurrent downloads: MAX_CONCURRENT_INPUT_DOWNLOADS bounds how many remote inputs a single request fetches at once, preventing socket/memory exhaustion and SSRF amplification from a request carrying many URLs.

Application limits do not cap the total request body

MAX_INPUT_FILE_SIZE applies to individual file inputs. It does not bound the overall request body: a large prompt sent as plain text (for example a huge messages array rather than a file) is still received and parsed in full before any application limit applies. Cap the total body size at the edge (below).

Edge body-size limit (recommended):

The overall request body should be capped at the network edge, before it is buffered by the application — this is the primary protection against memory exhaustion from oversized bodies.

  • AWS WAF: the mechanism for an ALB-fronted deployment. Add a rule with a SizeConstraintStatement on the request BODY and set oversize handling to block, so bodies larger than the inspectable size are rejected outright. The Terraform module's WAF (alb_waf_enabled=true) is the place to add this rule. AWS WAF inspects only a portion of the body (8 KB by default, up to 64 KB on ALB), so rely on the oversize-handling block action rather than attempting to inspect the full payload.
  • ALB: does not enforce a request body size limit on its own — pair it with AWS WAF for that control.
  • Amazon API Gateway: if you front the service with API Gateway instead of an ALB, note its hard 10 MB payload limit, which may be too small for large multimodal or long-context requests.

Sizing the edge limit is inherently hard for AI workloads

There is no universally correct body-size cap: legitimate prompts are large and keep growing as model context windows expand, and multimodal inputs (images, audio, documents) push request bodies higher still. Set the edge limit generously — large enough for your largest legitimate request, small enough to stop obvious abuse — and rely on the finer-grained controls below for cost and abuse management rather than a single tight byte limit.

Duration & throughput controls:

  • Output length: bound generated output with the API's max_tokens / max_output_tokens parameters; large client-supplied defaults can be constrained at your application or gateway layer.
  • Upstream timeout: AI_RESPONSE_TIMEOUT closes stalled model connections so a slow or hung generation does not hold resources indefinitely.
  • Rate limiting & concurrency: use WAF rate limiting (alb_waf_rate_limit, see Security Best Practices) to bound requests per client, complementing the per-request download concurrency cap above.

AWS Security Hub, GuardDuty & DNS Firewall Integration

Recommended: deploy via the Terraform module

The stdapi-ai Terraform module is built against the AWS Security Hub Foundational Security Best Practices (FSBP) standard and passes a large share of its applicable controls by default — no extra configuration required. A set of opt-in variables closes the remaining gaps for organizations with stricter compliance requirements, and the same dedicated VPC also enables native GuardDuty Runtime Monitoring and Route 53 Resolver DNS Firewall (see below). This makes the module the fastest path to a Security Hub-compliant, threat-monitored deployment; manual (non-Terraform) deployments must implement each control yourself.

The module composes three child modules — VPC, KMS, and ECS Fargate — plus its own ALB/WAF/S3/IAM resources. Each repository's README documents every relevant control — pass, fail, conditional (depends on your configuration), or not applicable — with remediation notes where relevant.

Module Key controls covered by default Extra options to improve compliance
stdapi-ai (root) S3 encryption/versioning/public-access block, least-privilege IAM, ALB access logging alb_waf_enabled, alb_certificate_arn / alb_domain_name (HTTPS redirect), deletion_protection
VPC EC2.2 (default security group lockdown), EC2.6 (flow logs, 365-day retention), EC2.15/21/48, IAM.24 compliance_vpc_endpoints_enabled, guardduty_vpc_endpoint_enabled, dns_firewall_enabled — dedicated VPC only, see below
KMS KMS.3 (deletion window), KMS.4 (automatic rotation, hardcoded on), KMS.5 (key policy) None — rotation and deletion window are hardcoded, not configurable
ECS Fargate ECS.*, EFS.1-8 (POSIX user enforcement, native backups), CloudWatch.15-17, Backup.1-5, IAM.1/24 mount_points_efs_backup_enable, mount_points[].efs_posix_user, alarms_enabled

All four modules also accept a tags variable to propagate custom resource tags — itself relevant to Security Hub controls IAM.24 and EC2.48 when a tagging policy is enforced on your account.

GuardDuty Runtime Monitoring

stdapi.ai's ECS tasks support Amazon GuardDuty Runtime Monitoring natively. Set guardduty_vpc_endpoint_enabled = true and the Terraform module creates the dedicated guardduty-data interface VPC endpoint the GuardDuty agent needs to report findings — keeping that traffic off the internet and off any NAT gateway. This isn't a Security Hub control, so it stays disabled by default; enable it whenever GuardDuty Runtime Monitoring is turned on for your account, since provisioning the endpoint through the module guarantees correct subnet placement.

Route 53 Resolver DNS Firewall

stdapi.ai fetches user-supplied URLs as multimodal content references (images, documents, audio) — see SSRF Protection for the application-level defense (IP-based blocklisting of loopback/link-local/private ranges). Route 53 Resolver DNS Firewall adds a network-layer control on top of that: set dns_firewall_enabled = true and the Terraform module creates a DNS Firewall rule group on the dedicated VPC that blocks or alerts on outbound DNS queries to known-malicious domains, closing the gap SSRF protection doesn't cover — a public IP address that resolves from a domain associated with malware, phishing, or botnet command-and-control.

  • dns_firewall_managed_domain_list_ids (default: null, which resolves automatically to the AWS Managed Aggregate Threat List ID for your deployment region — no AWS CLI call or extra IAM permissions needed) selects which managed domain lists to enforce; dns_firewall_action controls whether matches are BLOCKed, ALERTed on, or explicitly ALLOWed.
  • dns_firewall_advanced_enabled turns on DNS Firewall Advanced (additional AWS cost), which detects domain generation algorithm (DGA) and DNS tunneling activity — patterns a static domain list can't catch. Tune sensitivity with dns_firewall_advanced_confidence_threshold (LOW / MEDIUM / HIGH).

compliance_vpc_endpoints_enabled, guardduty_vpc_endpoint_enabled, and dns_firewall_enabled require the module's dedicated VPC

All three variables attach resources (interface endpoints, or a DNS Firewall rule group) to the VPC the Terraform module creates for you. They have no effect — dns_firewall_enabled cannot even be set to true and fails validation — if you instead pass your own subnet_ids to integrate with existing network infrastructure (see Integration with Existing Infrastructure). In that case the module never creates a VPC, and you are responsible for provisioning any interface endpoints or DNS Firewall rules your compliance or monitoring tooling needs (ECR, SSM, guardduty-data, Resolver rule groups, etc.) directly in your own VPC.


Prompt Injection

Prompt injection is an application-level security concern: just as AWS secures the database engine but customers are responsible for preventing SQL injection, AWS secures the Bedrock infrastructure but your application is responsible for preventing prompt injection. For full details, see the Amazon Bedrock prompt injection security documentation.

Key practices recommended by AWS:

  • Input validation — validate and sanitize all user input before passing it to the model; remove or escape special characters and enforce expected formats.
  • Secure coding — avoid string concatenation for building prompts; apply the principle of least privilege when granting access to resources.
  • System prompt scoping — when using system prompts, clearly define what the model can and cannot do. Newer models differentiate between system and user prompts — check the model provider's documentation for model-specific guidance.
  • Security testing — regularly test for prompt injection using penetration testing, static analysis, and dynamic application security testing (DAST).
  • Keep dependencies updated — monitor AWS security bulletins and keep the Bedrock SDK and libraries up to date.

Use Amazon Bedrock Guardrails

The most effective mitigation available within stdapi.ai is to configure an Amazon Bedrock Guardrail. Guardrails include a dedicated prompt attack detection layer and can be applied per request (via request headers) or globally for the entire server (via environment variables). See the Bedrock Guardrails configuration section for setup instructions, and Detect prompt attacks with Amazon Bedrock Guardrails for the upstream AWS documentation.

stdapi.ai does not include built-in prompt injection protection

stdapi.ai passes user prompts to the model as-is and does not apply any custom filtering or sanitization against prompt injection. Protecting against this risk is the responsibility of your application and infrastructure.


Security Best Practices

Network Isolation:

  • Deploy in private subnets — without direct internet access.
  • Restrict security group inbound rules — allow traffic to the stdapi.ai service only from the ALB or API Gateway security group on the application port.
  • Apply strict egress rules — limit outbound traffic to only the required AWS services (Bedrock, S3, etc.).
  • Use VPC endpoints — the Terraform module creates VPC endpoints for all required AWS services (vpc_endpoints_allowed=true by default), keeping all traffic to Bedrock, S3, SSM, CloudWatch Logs, and AI services off the public internet with no NAT Gateway required.

WAF & Rate Limiting:

  • Enable AWS WAF — the Terraform module supports WAF via alb_waf_enabled=true. When enabled, it activates the AWS Managed Core Rule Set (SQLi, XSS, common exploits), Linux Rule Set, and IP Reputation List (known malicious IPs).
  • Configure rate limiting — set alb_waf_rate_limit to limit the number of requests per IP per 5-minute window, protecting against abuse and quota exhaustion.
  • Block anonymous IPs — set alb_waf_block_anonymous_ips=true to block known VPNs, open proxies, and Tor exit nodes.
  • Cap request body size — add an AWS WAF SizeConstraintStatement on the request body (oversize handling set to block) to reject oversized bodies at the edge; see Request Size & Resource Limits.

Secrets Management:

  • Use SSM Parameter Store or Secrets Manager — always use encrypted secret storage for production API keys; never pass keys as plain environment variables.
  • Rotate API keys — Secrets Manager supports automated secret rotation via Lambda; note that stdapi.ai reads the key once at startup, so a container restart is required after rotation.
  • Share the Realtime signing key — the ephemeral client secrets minted for browser clients are signed, not stored, so every instance must sign with the same key. It is derived from the deployment's API key by default; a deployment running with no API key at all falls back to a per-process random key and must set REALTIME_CLIENT_SECRET_KEY instead. Treat it as a secret and store it the same way: anyone holding it can mint a client secret this deployment accepts.

Eliminate key rotation with AWS native auth

When using API key authentication, rotation requires updating SSM/Secrets Manager and performing a rolling ECS task replacement. This is predictable but adds a deployment step.

To avoid key rotation entirely, use Amazon Cognito user pool tokens, or OIDC / Cognito / IAM Identity Center via the ALB — there are no API keys to rotate, and access is granted and withdrawn in your identity provider without touching the deployment. With tokens validated by stdapi.ai, withdrawing access takes effect when the caller's current token expires: keep the app client's access-token validity short.

IAM & Access Control:

  • Apply least-privilege IAM permissions — ensure the ECS Task Role has only the required permissions needed for operation.

Terraform Module

The Terraform module enforces security defaults out of the box: least-privilege IAM roles scoped to the minimum permissions required by each ECS task, KMS-encrypted SSM parameters, log groups, and S3 buckets, and network isolation with ECS tasks in private subnets and Security Group rules that allow only ALB-to-service traffic.


Encryption in Transit

To ensure data protection during transmission, stdapi.ai supports several encryption-in-transit configurations.

Mode TLS scope Config required Typical use case
Standard Entry point only None Default for most AWS deployments
End-to-End Entry point → container GRANIAN_SSL_CERTIFICATE + GRANIAN_SSL_KEYFILE Compliance-mandated full-path encryption
mTLS at Container Entry point → container, mutual Above + GRANIAN_SSL_CA + GRANIAN_SSL_CLIENT_VERIFY Mutual authentication between ALB and container
mTLS at Entry Point Client → entry point, mutual ALB or API Gateway configuration Client certificate enforcement before traffic reaches the application

Standard

This is the default configuration when no additional encryption settings are applied. TLS is terminated at the entry point (ALB or API Gateway). Traffic between the entry point and the stdapi.ai service travels over HTTP within the private VPC, secured by mandatory Security Group rules. This is the most common configuration for AWS deployments.

Terraform Module

The Terraform module configures the ALB with the ELBSecurityPolicy-TLS13-1-2-Res-PQ-2025-09 security policy by default (alb_ssl_policy), enforcing a minimum of TLS 1.2 and enabling TLS 1.3 with post-quantum hybrid key exchange. See Data Sovereignty & Compliance for details on ALB TLS configuration and post-quantum key exchange support.

End-to-End

If your compliance requirements mandate encryption all the way to the container, you can enable TLS between the load balancer and the container.

  • Native Support: The provided container images run Granian natively. TLS can be enabled directly by setting the following Granian environment variables:
    • GRANIAN_SSL_CERTIFICATE: Path to the SSL certificate file (e.g., /etc/ssl/certs/server.crt).
    • GRANIAN_SSL_KEYFILE: Path to the SSL private key file (PKCS#8 format only, e.g., /etc/ssl/private/server.key).
    • GRANIAN_SSL_KEYFILE_PASSWORD: (Optional) Password for the private key file.
    • GRANIAN_SSL_PROTOCOL_MIN: (Optional) Minimum supported TLS version — tls1.2 or tls1.3 (defaults to tls1.3).
  • Mutual TLS — Container: As an advanced form of end-to-end encryption, you can also require the upstream caller (e.g., the ALB) to present a certificate to the container.
    • GRANIAN_SSL_CA: Path to the CA certificate bundle used to verify client certificates.
    • GRANIAN_SSL_CLIENT_VERIFY: Set to true to enable client certificate verification.
  • ECS Certificate Mounting: When running on AWS ECS, certificate and key files must be mounted into the container. This is typically achieved using Amazon EFS volumes or by baking the certificates into a custom image (not recommended for secrets).
  • Certificate Management: Use ACM Private CA or a self-signed certificate for the internal service. Note that the ALB does not validate the backend certificate by default; it only ensures the connection is encrypted.
  • Sidecar Proxy: Alternatively, you can run a sidecar container (like Nginx or Envoy) alongside stdapi.ai to handle TLS termination locally.

Terraform Module Support

The provided Terraform module does not currently support automatic configuration of end-to-end encryption. You will need to manually configure the ALB target group for HTTPS and manage certificate provisioning.

Mutual TLS — Client to Entry Point

For environments requiring mutual authentication between the final client and the entry point:

  • Offloaded mTLS: mTLS can be enabled at the ALB (using Application Load Balancer Mutual TLS) or API Gateway (using Mutual TLS authentication). In this setup, the entry point validates the client certificate before forwarding the request.
  • Service Mesh / Sidecar: You can implement mTLS between internal services using a Service Mesh (like AWS App Mesh) or by using a sidecar proxy (like Envoy) that manages certificate rotation and verification.

Next Steps