Ollama Rate Limits
Local Ollama (http://localhost:11434) has no rate limits or authentication. Ollama Cloud enforces tier-based concurrency (Free, Pro=3, Max=10 concurrent cloud models) and weekly GPU-time quotas rather than per-second request ceilings. Cloud quotas reset on 5-hour session and 7-day weekly cycles. Specific TPS / RPM ceilings are not publicly documented.
Ollama Rate Limits is the machine-readable rate-limit profile for Ollama on the APIs.io network, conforming to the API Commons Rate Limits specification.
It captures 5 rate-limit definitions, measuring requests_per_second, concurrent_requests, and gpu_time.
The profile also includes 4 backoff/retry policies defined and response codes documented for unauthorized, throttled, and serviceUnavailable.
Tagged areas include Artificial Intelligence, Large Language Models, Models, and Rate Limiting.
Limits
Policies
Sources
Work with this as data
Every rate limit here is available over the APIs.io API and to AI agents over MCP.
MCP server
One button, every client — Claude, Cursor, VS Code and the rest.
https://apis.io/mcp
Tools for rate limits
4 MCP tools reach this
find_rate_limitsBrowse and filter every rate limit in the catalog.apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.resolveTurn a domain, URL or GitHub org into the provider it belongs to.find_cohortsEvery scored population of providers in the catalog.
Call it yourself
curl for this page
curl "https://apis.io/api/v1/rate-limits/ollama-rate-limits"
curl "https://apis.io/api/v1/rate-limits?limit=25"
Discovery needs no key. Ratings and market analysis are Pro.