Scalable Inference Serving Rate Limits
Scalable Inference Serving does not publish public API rate limits reachable in this reconciliation pass; limits are governed by the customer's commercial / partner agreement. Consumers should honor 429 / 503 responses where they appear and follow standard exponential-backoff guidance.
Scalable Inference Serving Rate Limits is the machine-readable rate-limit profile for Scalable Inference Serving on the APIs.io network, conforming to the API Commons Rate Limits specification.
It captures 1 rate-limit definition, measuring varies.
The profile also includes 2 backoff/retry policies defined and response codes documented for throttled and serviceUnavailable.
Tagged areas include AI, Inference, Kubernetes, and Rate Limiting.
Limits
Policies
Sources
- https://kserve.github.io/website/
- https://docs.bentoml.com/
- https://docs.vllm.ai/
- https://github.com/triton-inference-server/server
Work with this as data
Every rate limit here is available over the APIs.io API and to AI agents over MCP.