Amazon Nova · Rate Limits

Amazon Nova Rate Limits

Published Amazon Bedrock service quotas that govern Amazon Nova model invocation, harvested from the AWS General Reference quota tables on 2026-09-01. This file REPLACES an API Evangelist scaffold dated 2026-05-04 that asserted invented per-key request quotas and X-RateLimit-* headers; none of those existed. Amazon Nova is not rate limited per API key at all — it is quota'd per AWS account, per region, per MODEL, in requests-per-minute and tokens-per-minute.

Amazon Nova Rate Limits is the machine-readable rate-limit profile for Amazon Nova on the APIs.io network, conforming to the API Commons Rate Limits specification.

It captures 20 rate-limit definitions, measuring requests_per_minute, tokens_per_minute, and concurrent_requests.

The profile also includes 3 backoff/retry policies defined and response codes documented for throttled, quotaExceeded, and serviceUnavailable.

Tagged areas include Foundation Models, Rate Limiting, Quotas, and Throttling.

20 Limits Throttle: 429 Quota: 400
Foundation ModelsRate LimitingQuotasThrottling

Limits

On-demand model inference requests per minute for Amazon Nova Micro account-region-model
requests_per_minute
2000
On-demand model inference tokens per minute for Amazon Nova Micro account-region-model
tokens_per_minute
4000000
On-demand model inference requests per minute for Amazon Nova Lite account-region-model
requests_per_minute
2000
On-demand model inference tokens per minute for Amazon Nova Lite account-region-model
tokens_per_minute
4000000
On-demand model inference requests per minute for Amazon Nova Pro account-region-model
requests_per_minute
250
On-demand model inference tokens per minute for Amazon Nova Pro account-region-model
tokens_per_minute
1000000
On-Demand, latency-optimized model inference requests per minute for Amazon Nova Pro V1 account-region-model
requests_per_minute
10
On-Demand, latency-optimized model inference tokens per minute for Amazon Nova Pro V1 account-region-model
tokens_per_minute
40000
Cross-region model inference requests per minute for Amazon Nova Premier V1 account-region-model
requests_per_minute
500
Cross-region model inference tokens per minute for Amazon Nova Premier V1 account-region-model
tokens_per_minute
2000000
On-demand model inference requests per minute for Amazon Nova Canvas account-region-model
requests_per_minute
100
On-demand InvokeModel concurrent requests for Amazon Nova Reel 1.0 account-region-model
concurrent_requests
10
On-demand InvokeModel concurrent requests for Amazon Nova Reel 1.1 account-region-model
concurrent_requests
3
On-demand InvokeModel concurrent requests for Amazon Nova Sonic account-region-model
concurrent_requests
20
Cross-region model inference requests per minute for Amazon Nova 2 Lite account-region-model
requests_per_minute
2000
Cross-region model inference tokens per minute for Amazon Nova 2 Lite account-region-model
tokens_per_minute
8000000
Cross-region model inference requests per minute for Amazon Nova 2 Omni account-region-model
requests_per_minute
2000
Cross-region model inference tokens per minute for Amazon Nova 2 Omni account-region-model
tokens_per_minute
8000000
Cross-region model inference requests per minute for Amazon Nova 2 Pro Preview account-region-model
requests_per_minute
100
Cross-region model inference tokens per minute for Amazon Nova 2 Pro Preview account-region-model
tokens_per_minute
1000000

Policies

Backoff Strategy
No Retry-After is returned, so a caller must implement exponential backoff with jitter on its own. ModelNotReadyException is the only error Amazon marks retryable in the service model; AWS SDKs retry it automatically.
Quota Discovery
Current and remaining quota is not observable from a response. Read it from the Service Quotas console or API (service code `bedrock`) and from CloudWatch metrics.
Quota Increase
Token-per-minute quotas are generally adjustable via Service Quotas; on-demand requests-per-minute quotas for most Nova models are marked not adjustable.

Work with this as data

Every rate limit here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for rate limits

4 MCP tools reach this
  • find_rate_limitsBrowse and filter every rate limit in the catalog.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This rate limit
curl "https://apis.io/api/v1/rate-limits/amazon-nova-rate-limits"
All rate limits
curl "https://apis.io/api/v1/rate-limits?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.