CentML · Rate Limits

Centml Rate Limits

CentML's serverless inference API enforces per-account rate limits expressed as requests per minute and tokens per minute, which vary by model and account tier. Dedicated deployments are governed by the capacity of the provisioned GPU hardware and configured autoscaling (min/max replicas) rather than shared account-level token limits. Specific per-model limit values are not reconciled in this artifact.

Centml Rate Limits is the machine-readable rate-limit profile for CentML on the APIs.io network, conforming to the API Commons Rate Limits specification.

It captures 4 rate-limit definitions, measuring requests, tokens, and replicas.

The profile also includes 2 backoff/retry policies defined and response codes documented for throttled.

Tagged areas include AI, LLM, Inference, Serverless, and GPU.

4 Limits Throttle: 429
AILLMInferenceServerlessGPURate LimitingQuotasThrottling

Limits

Requests Per Minute (RPM) account
requests
see provider documentation
Per-model RPM on serverless endpoints, varies by tier and model.
Tokens Per Minute (TPM) account
tokens
see provider documentation
Per-model TPM on serverless endpoints, varies by tier and model.
Concurrent Requests account
requests
see provider documentation
Concurrency permitted against serverless endpoints, varies by tier.
Dedicated Replica Capacity deployment
replicas
configured via min/max replicas
Dedicated deployment throughput is bounded by provisioned GPU hardware and autoscaling settings, not shared account token limits.

Policies

Tiered Limits
Limits raise as accounts move from free / trial to paid usage and via Enterprise agreements.
Backoff Strategy
Clients should implement exponential backoff with jitter and honor Retry-After on 429 responses.

Sources

Work with this as data

Every rate limit here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for rate limits

4 MCP tools reach this
  • find_rate_limitsBrowse and filter every rate limit in the catalog.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This rate limit
curl "https://apis.io/api/v1/rate-limits/centml-rate-limits"
All rate limits
curl "https://apis.io/api/v1/rate-limits?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.