Featherless AI · Rate Limits

Featherless Rate Limits

Featherless AI does not meter tokens. Instead each subscription plan grants a fixed number of concurrent connections (concurrency limit), and tokens are unlimited within that concurrency. Smaller models allow higher effective throughput than larger ones. The number of in-flight requests an account may run is therefore the primary limit; additional requests beyond the plan concurrency are rejected until a slot frees. Live concurrency state can be observed via the account concurrency stream.

Featherless Rate Limits is the machine-readable rate-limit profile for Featherless AI on the APIs.io network, conforming to the API Commons Rate Limits specification.

It captures 3 rate-limit definitions, measuring connections, tokens, and parameters.

The profile also includes 2 backoff/retry policies defined and response codes documented for throttled.

Tagged areas include AI, LLM, Inference, Serverless, and Open Models.

3 Limits Throttle: 429
AILLMInferenceServerlessOpen ModelsRate LimitingQuotasThrottling

Limits

Concurrent Connections account
connections
see plan (Basic 2, Premium 4, agent/business 8 per unit)
Maximum simultaneous in-flight inference requests per account/plan.
Tokens account
tokens
unlimited within plan concurrency
Featherless bills a flat monthly subscription; tokens are not metered.
Model Size account
parameters
see plan (Basic up to 15B; Premium any size)
Larger models reduce effective per-connection throughput.

Policies

Concurrency-Based Limiting
Throughput is governed by the number of concurrent connections in the plan rather than by token quotas; excess concurrent requests are rejected with HTTP 429 until a slot is available.
Backoff Strategy
Clients should implement exponential backoff with jitter and honor Retry-After when concurrency is saturated.

Sources

Work with this as data

Every rate limit here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for rate limits

4 MCP tools reach this
  • find_rate_limitsBrowse and filter every rate limit in the catalog.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This rate limit
curl "https://apis.io/api/v1/rate-limits/featherless-rate-limits"
All rate limits
curl "https://apis.io/api/v1/rate-limits?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.