Modal Labs Rate Limits
Modal governs usage through concurrency limits rather than classic requests-per-minute API throttling, because the developer surface is an SDK/gRPC control plane and user-deployed containers, not a metered REST API. The key limits are concurrent containers and concurrent GPUs, which scale by subscription tier. User-deployed web endpoints (*.modal.run) scale by container concurrency; any request-level rate limiting on those endpoints is defined by the developer's own function code.
Modal Labs Rate Limits is the machine-readable rate-limit profile for Modal on the APIs.io network, conforming to the API Commons Rate Limits specification.
It captures 6 rate-limit definitions, measuring containers, gpus, and requests.
The profile also includes 3 backoff/retry policies defined and response codes documented for throttled.
Tagged areas include Serverless, Compute, GPU, AI Infrastructure, and Rate Limiting.
Limits
Policies
Sources
- https://modal.com/pricing
- https://modal.com/docs/guide/concurrent-inputs
- https://modal.com/docs/guide/scale
Work with this as data
Every rate limit here is available over the APIs.io API and to AI agents over MCP.
MCP server
One button, every client — Claude, Cursor, VS Code and the rest.
https://apis.io/mcp
Tools for rate limits
4 MCP tools reach this
find_rate_limitsBrowse and filter every rate limit in the catalog.apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.resolveTurn a domain, URL or GitHub org into the provider it belongs to.find_cohortsEvery scored population of providers in the catalog.
Call it yourself
curl for this page
curl "https://apis.io/api/v1/rate-limits/modal-labs-rate-limits"
curl "https://apis.io/api/v1/rate-limits?limit=25"
Discovery needs no key. Ratings and market analysis are Pro.