Cost Management¶
Running stdapi.ai bills into three independent buckets:
| Cost | Billed by | Scales with | Covered by AWS credits |
|---|---|---|---|
| AI service usage | AWS (Amazon Bedrock, Polly, Transcribe, …) or AWS Marketplace, per model | Requests and tokens | Only the AWS-billed models |
| Gateway infrastructure | AWS (ECS Fargate, ALB, S3, CloudWatch, …) | Running containers | Yes |
| Gateway license | AWS Marketplace (Enterprise Edition only) | Container-hours | No — Marketplace is never credited |
AI service usage dominates any normal deployment. The gateway runs on a small ARM64 container, so its infrastructure and license form a low fixed floor that model spend passes quickly. Read the AI-cost sections first; Gateway Cost matters mainly at low traffic or under a hard cost constraint.
stdapi.ai adds no markup: model usage is billed to you by AWS at the same rate as calling Bedrock directly.
Knowing the Price Before You Call¶
The Model Pricing API (GET /model_pricing) exposes the exact AWS unit prices stdapi.ai resolved for each model, from the same catalog it uses to compute costs. Use it for cost-aware model selection — comparing candidates before routing traffic, rather than discovering the rate afterwards in the logs.
Marketplace-Billed vs AWS-Billed Models¶
A separate AWS Marketplace line on your bill for Bedrock usage is expected, not an error. AWS states it plainly:
Amazon Bedrock FAQ — Why do I see a billing entry for AWS Marketplace for my usage of AWS Bedrock?
"Customers will see an AWS Marketplace bill for certain Bedrock serverless models and Bedrock Marketplace models. This is because these models are sold by third party providers as Third-Party Content, as described in the AWS service terms section 50.12."
Why It Matters: AWS Credits¶
This is the practical consequence, and it is easy to get wrong when planning a budget:
Promotional credits do not apply to AWS Marketplace charges
The AWS Promotional Credit Terms exclude AWS Marketplace from eligible services: "Promotional Credit will not be applied to any fees or charges for Amazon Mechanical Turk, AWS Managed Services, Ineligible AWS Support, AWS Marketplace, … (collectively, 'Ineligible Services')."
A model billed through AWS Marketplace therefore consumes real money even when your account holds AWS credits, while the same workload on an AWS-billed model would draw them down. If you are running on credits, promotional funding, or an activate/startup programme, model choice changes your actual outlay — not just your reported spend.
The same applies to committed-spend agreements: check whether your EDP or Private Pricing Agreement counts Marketplace spend before assuming a discount applies.
Which Models Are Which¶
The split follows the model provider, not the API you call. Per AWS, models from Amazon, DeepSeek, Mistral AI, Meta, Qwen and OpenAI are not sold through AWS Marketplace, so they bill as ordinary Amazon Bedrock usage. Other providers — Anthropic among them — are third-party listings billed through Marketplace.
Verify rather than assume
That split changes as AWS onboards providers, so treat the list above as a starting point, not a contract. Two reliable checks:
- A model that needs
aws-marketplace:Subscribeon first use is a Marketplace listing. AWS documents the permission error this produces. - In the AWS Price List, Marketplace listings carry an "(Amazon Bedrock Edition)" product name. This is the same signal stdapi.ai's price catalog uses to parse them.
What stdapi.ai Does About It¶
- Both paths are priced. Cost tracking ingests Marketplace listings and native Bedrock rows alike, so a request log entry carries a cost regardless of how AWS bills it.
- Subscriptions are handled automatically. AWS creates the Marketplace subscription on first invocation; stdapi.ai performs that flow when
AWS_BEDROCK_MARKETPLACE_AUTO_SUBSCRIBEis enabled (the default). It requiresaws-marketplace:Subscribeandaws-marketplace:ViewSubscriptions— see IAM Permissions. Without them, the first call to a third-party model fails withAccessDeniedException. - Cost tracking does not separate the two lines. Request logs report what a call cost, not which AWS invoice section it lands on. Use Cost Explorer, grouped by service, to see the Marketplace split.
Cost Tracking (Real-Time AWS Pricing)¶
When COST_TRACKING=true is enabled, stdapi.ai computes real-time costs from live AWS pricing. Costs are attributed to the actual region where each request was served.
How It Works¶
The loaded catalog is also queryable through the Model Pricing API (GET /model_pricing) for cost-aware model selection.
- Price Catalog: At startup, stdapi.ai fetches the AWS Price List API for all configured regions and services in a background task, then caches in memory. Server readiness never waits on it: requests served before the load completes simply record usage without a cost. Failed loads are retried with exponential backoff (1–15 min), and each attempt's outcome is logged as a
backgroundevent namedprice_catalog_load - On-Demand Refresh: The catalog is refreshed whenever a newly available Bedrock model is discovered with no catalog entry yet — not on a proactive schedule. If that refresh's AWS call fails, the error fails the request that triggered it
- Per-Request Computation: For each request, costs are computed by multiplying billed quantities by the resolved unit price
- Built-in Defaults: A few models are priced only on the AWS pricing page, not in the Price List API (e.g. the Stability AI Image Services) — their page rates ship built in, used only when AWS publishes no row for the model;
COST_PRICE_OVERRIDESstill takes precedence - Fallback on a Missing Price: Once the catalog has been fetched, if a specific model/dimension has no resolvable price in it, the cost field is omitted for that entry rather than blocking the request, and the request log carries a
warning-levelerror_detailnaming the model and unpriced dimensions (a hint to supply the missing price viaCOST_PRICE_OVERRIDES)
Pricing is an estimate, not a bill
stdapi.ai resolves prices from AWS's own Price List API and does its best to match every request to the right unit price — including tier, cache TTL, cross-region routing, region fallback, and image resolution/quality where applicable. This is still a best-effort approximation, not a guarantee: AWS's Price List API doesn't reliably map a Bedrock model ID to its own pricing rows, some pricing dimensions aren't modeled at all (see Known Limitations), and fallbacks (regional, tier) substitute a nearby price when the exact one isn't published. For billing-critical use, always reconcile against AWS Cost Explorer or your actual invoice.
Configuration¶
| Setting | Default | Description |
|---|---|---|
COST_TRACKING |
false |
Enable/disable real-time cost computation (needs pricing:GetProducts) |
COST_PRICE_OVERRIDES |
{} |
JSON map for operator-supplied prices for models not in AWS catalog |
Request Log Format¶
Each usage entry includes cost and currency when resolved:
{
"service": "bedrock-runtime",
"model": "anthropic.claude-3-5-sonnet-20241022-v2:0",
"operation": "/v1/chat/completions",
"region": "us-east-1",
"input_tokens": 1500,
"output_tokens": 450,
"cost": "0.004575",
"currency": "USD"
}
Costs are exact plain-decimal strings (never floats), so no precision is lost and no exponent or trailing zeros ever appear. The request-level total is also logged:
{
"cost": {
"USD": "0.012345"
}
}
EMF Cost Metric¶
When CloudWatch metrics are enabled, a Cost metric (unit: None) is emitted under the ["Model", "Currency"] dimension set — a separate directive from the quantity metrics, which stay under ["Model"]. A single directive spanning both sets would also publish Cost bare-by-Model, silently summing across currencies. Query Cost with both Model and Currency dimensions:
{
"_aws": {
"CloudWatchMetrics": [
{
"Namespace": "stdapi",
"Dimensions": [["Model"]],
"Metrics": [{"Name": "InputTokens", "Unit": "Count"}]
},
{
"Namespace": "stdapi",
"Dimensions": [["Model", "Currency"]],
"Metrics": [{"Name": "Cost", "Unit": "None"}]
}
]
},
"Model": "anthropic.claude-3-5-sonnet-20241022-v2:0",
"Currency": "USD",
"Cost": 0.004575,
"InputTokens": 1500
}
Regional Price Fallback¶
Some models — mostly older/deprecated ones — aren't published in every region's Price List (e.g. priced in us-east-1 but not any EU region). A region with no price for a given model/dimension/tier always borrows one from a nearby region instead of omitting the cost:
- Prefers another region in the same geography (
eu-west-3tries othereu-*regions first) - Falls back to
us-east-1,eu-west-1, orus-west-2— the regions always fetched regardless of your configured Bedrock regions - If neither is available, the cost is omitted as usual
This is a substitute price, not the actual published price for that region.
Multi-Currency Support¶
stdapi.ai detects currency from the AWS partition: - Standard AWS: USD - AWS European Sovereign Cloud (EUSC): EUR - AWS US GovCloud: USD - AWS China: CNY
Costs are never summed across currencies — this safety behavior is always on, regardless of settings. It matters when a single request's billed dimensions resolve to different currencies, which can happen with regional fallback crossing a partition boundary (e.g. a EUSC deployment falling back to a standard-AWS-priced region). A usage entry that spans more than one currency reports a costs map (instead of cost/currency) with every currency's own amount:
{
"service": "bedrock-runtime",
"model": "anthropic.claude-3-5-sonnet-20241022-v2:0",
"input_tokens": 1500,
"output_tokens": 450,
"costs": {
"EUR": "0.0021",
"USD": "0.0028"
}
}
When only one currency is involved, the entry's cost/currency reports that currency directly instead. The request-level cost total still aggregates cleanly per currency either way.
Routing-Tier Pricing¶
AWS prices some models differently per serving profile: the cross-region "global" routing profile is lower than the plain/regional rate for some marketplace-listed models (confirmed live: Claude Sonnet 4.5 input tokens at $3.30/M regional vs $3.00/M global), while latency-optimized serving (requested via the X-Amzn-Bedrock-PerformanceConfig-Latency: optimized header) is higher. stdapi.ai tracks the profile that served each request ("routing": "global" or "latency") and prices it at the matching rate, falling back to the regional rate when AWS publishes no distinct one.
Long-Context Pricing¶
Some 1M-context-capable models (currently Claude Sonnet 4, via the context-1m anthropic-beta flag) are billed by AWS at a higher rate — roughly double for input-side tokens — when a call's prompt (input + cache read/write tokens) exceeds 200K tokens. stdapi.ai detects this per model call and prices the whole call at the published long-context rate, reporting it with "context": "long" in the usage entry. When AWS publishes no long-context rate for a model, the standard rate is used as the best available estimate.
Built-in Tool Pricing¶
Amazon Nova's built-in grounding tool (web_search → nova_grounding, supported by Nova 2 and Nova Premier) is billed by AWS per grounding request, on top of token usage. stdapi.ai counts grounding invocations in the model response and reports them as grounding_requests, priced from the model's published per-request rate.
Image Pricing Granularity¶
Some image-generation models (currently: Amazon Titan Image Generator G1/V2 and Nova Canvas) are priced by AWS per resolution/quality combination, not a flat per-image rate. stdapi.ai automatically prices these per-image, based on the actual requested size and quality, falling back to a flat per-image rate for models where this isn't wired up yet (e.g. Stability). No configuration needed — this is purely additive precision with no accuracy trade-off.
Media Input Pricing¶
Multimodal embedding inputs are billed by AWS per media unit on top of (or instead of) tokens: per input image (with a distinct "document image" rate where offered) and per second of audio/video. stdapi.ai records image counts directly and audio/video durations from the AWS-reported segment timings of the asynchronous (segmented) processing path, reported as input_images/input_seconds with their *_by_spec breakdowns. Rerank queries are recorded as search_units (one per query).
Known Limitations¶
- Speech-modality tokens (speech-to-speech models): AWS bills speech tokens at a higher rate than text tokens, but reported token usage doesn't distinguish modalities — all tokens are priced at the text rate, which underestimates speech-heavy calls.
- Asynchronous (segmented) embeddings: AWS reports no token usage for this processing path (used automatically for large inputs), so segmented text embeddings report no token cost; audio/video durations are recovered from the AWS-reported segment timings and billed.
- Synchronous audio/video embedding inputs: media duration is only available from AWS on the segmented (asynchronous) processing path — small audio/video inputs processed synchronously report no duration and no per-second cost. No estimate is substituted (this app only reports AWS-confirmed real usage).
- Client disconnect during streaming: chat (Converse) streams drain the already-open Bedrock stream on disconnect to still capture the trailing usage event. For InvokeModel streams and image generation jobs, AWS bills the input tokens and everything generated up to the cancellation, but no final usage report arrives, so nothing is recorded for that call. No estimate is substituted.
- Rerank queries with more than 100 documents: AWS bills one search unit per 100 documents; the document count isn't visible at recording time, so one unit per query is recorded.
- Reserved capacity pricing: if a request explicitly asks for AWS's Reserved Capacity service tier, its cost is computed at the standard on-demand rate instead — Reserved Capacity uses a separate monthly-commitment pricing model this app doesn't ingest. Avoid relying on this app's cost figures for Reserved Capacity workloads.
- Some very new or region-specific models may have no published price anywhere yet — AWS publishes pricing after model availability, sometimes with a delay.
- AWS GovCloud: the Price List API has no GovCloud endpoint, so catalog prices cannot be fetched there — usage is still recorded, a startup warning is emitted, and only
cost_price_overridesentries produce costs.
Override Map for Missing Models¶
Some models (e.g., mid-tier Claude) may not appear in the AWS Price List API. Use cost_price_overrides to fill gaps (Bedrock models only — Polly/Transcribe/Translate/Comprehend prices always come from the catalog):
export COST_PRICE_OVERRIDES='{"anthropic.claude-3-5-sonnet-20241022":{"input_tokens":0.000003,"output_tokens":0.000015}}'
Prices are per one unit (token, character, second) in your partition's currency.
IAM Permissions¶
Ensure your IAM role includes pricing read access (the Price List API serves identical data from its three commercial endpoints, and sovereign partitions such as EUSC host their own; the nearest one is selected automatically from your configured Bedrock regions):
{
"Effect": "Allow",
"Action": ["pricing:GetProducts"],
"Resource": "*"
}
AWS Cost Attribution¶
Cost tracking prices each request as it happens. AWS-side attribution answers a different question — whose spend lands on the AWS bill, in Cost Explorer and CUR 2.0. The two are complementary: request logs for granular analysis, AWS attribution for invoicing and chargeback.
| Dimension | Mechanism | Reported in | Setup |
|---|---|---|---|
| Service / gateway | IAM principal attribution — AWS captures the caller identity | Cost Explorer, CUR 2.0 | Automatic; tag the execution role for finer breakdowns |
| Application / workload | Bedrock Project/Workspace (Bedrock Mantle models) | Cost Explorer, CUR 2.0 | Set AWS_BEDROCK_MANTLE_PROJECT |
| End user | stdapi-ai.user_id request metadata and job tags |
stdapi.ai logs, Bedrock invocation logs | Clients send user (OpenAI) or metadata.user_id (Anthropic) |
IAM Principal Attribution¶
Amazon Bedrock records the IAM identity behind every bedrock-runtime inference call and forwards it to Cost Explorer and CUR 2.0. Nothing to configure in stdapi.ai, and no change to client requests.
stdapi.ai calls AWS with a single execution role — the ECS task role — so all Bedrock spend is attributed to that one identity, the standard LLM-gateway pattern AWS documents. To break that total down, tag the execution role (for example team or cost-center), then activate those keys in the AWS Billing console under Cost allocation tags, filtering by type IAM principal.
Attribution grain
IAM principal attribution aggregates per usage type per day — it never yields a per-request cost. Use the request logs for that. It also covers bedrock-runtime only: Bedrock Mantle requests are attributed with Projects instead.
Per-User Attribution¶
Because every call shares the gateway's execution role, AWS itself cannot split Cost Explorer costs per end user. stdapi.ai attributes users at the request level instead: the client-supplied identifier is recorded as stdapi-ai.user_id in the request log — alongside that request's computed cost — and forwarded to Bedrock as request metadata, so per-user spend is aggregated from the logs rather than from the AWS bill.
Gateway Cost¶
Secondary for most deployments — read this section if you run under a hard cost constraint or at low traffic, where the gateway's fixed floor is a visible share of the bill.
Infrastructure¶
The gateway is a single small container. The Terraform module defaults to ARM64, 0.25 vCPU and 512 MiB — the smallest Fargate size — which keeps this bucket cheap, though it is still a floor you pay whether or not any request arrives.
The Terraform module provisions:
| Component | Notes |
|---|---|
| ECS Fargate service | The dominant infrastructure cost; 0.25 vCPU / 512 MiB ARM64 by default, times autoscaling_min_capacity |
| Application Load Balancer | Hourly rate plus LCU; skipped when you attach your own |
| S3 buckets (regional) | Input files, generated media; see AWS_S3_VIDEOS_EXPIRES_AFTER |
| CloudWatch logs, metrics, alarms | Grows with log verbosity and EMF metric volume |
| KMS, IAM, networking | Keys and roles; NAT/VPC endpoints only when the module creates the network |
| WAF (optional) | Per-rule and per-request charges when alb_waf_enabled = true |
Keeping It Low¶
-
Stop the service outside business hours.
autoscaling_schedule_stopandautoscaling_schedule_starttake cron expressions, so the cluster can run weekdays only and cost nothing at night — the single largest saving for an internal workload.autoscaling_schedule_stop = "cron(0 19 ? * MON-FRI *)" autoscaling_schedule_start = "cron(0 8 ? * MON-FRI *)" -
Reuse existing network infrastructure. Passing
subnet_idsandsecurity_group_iddeploys into your VPC and creates no additional NAT gateways or load balancers — see Integration with Existing Infrastructure. On a small deployment these are often larger than the gateway itself. - Scale to your real floor.
autoscaling_min_capacitysets how many containers run at idle. Because the Enterprise license is billed per container-hour, this number drives both infrastructure and license cost. - Use Fargate Spot and trim log retention, as the Cost-Optimized Deployment example does.
- Trim observability volume. Raising
LOG_LEVELand disabling request/response payload logging cut CloudWatch ingestion — see Controlling Log Verbosity. Note that EMF metric lines are not affected byLOG_LEVEL: they are written to stdout on every request wheneverCLOUDWATCH_METRICSis on, so disabling that setting is the only way to remove them.
Estimating before you deploy
stdapi.ai publishes no infrastructure estimate: the total depends on your region, replica count, schedule and whether the module creates networking. Price the component list above with the AWS Pricing Calculator for your own configuration.
License¶
stdapi.ai is dual-licensed, and the edition you run decides whether the gateway itself costs anything.
| Edition | License | Cost |
|---|---|---|
| Community | AGPL-3.0-or-later | Free — you must share your modifications under the same license |
| Enterprise | Commercial, via AWS Marketplace | $0.10 per container-hour, pay-as-you-go, no upfront or minimum |
The Enterprise Edition includes a 14-day free trial, and a private offer brings the rate to $0.09 per container-hour. There is no per-request or per-token component: the license tracks running containers only, so cost follows autoscaling_min_capacity/max_capacity rather than traffic.
The license is itself a Marketplace charge
Being an AWS Marketplace subscription, the Enterprise license appears on the Marketplace section of your bill and — like Marketplace-billed models — cannot be paid with AWS promotional credits.
Next Steps¶
- Logging & Monitoring — Usage metrics fields, CloudWatch EMF, and log verbosity
- Model Pricing API — Query AWS unit prices per model for cost-aware selection
- Configuration Reference —
COST_TRACKING,COST_PRICE_OVERRIDESand IAM permissions - Licensing — AGPL vs commercial license and AWS Marketplace subscription