NVIDIA NIM · Rate Limits

Nvidia Nim Rate Limits

Nvidia Nim Rate Limits is the machine-readable rate-limit profile for NVIDIA NIM on the APIs.io network, conforming to the API Commons Rate Limits specification.

It captures 6 rate-limit definitions.

The profile also includes 7 backoff/retry policies defined.

Tagged areas include Artificial Intelligence, Inference, Microservices, LLM, and Foundation Models.

6 Limits
Artificial IntelligenceInferenceMicroservicesLLMFoundation ModelsGPUKubernetesNVIDIAOpenAI-Compatible

Policies

Hosted Developer RPM
Free developer tier soft rate limit on /v1/chat/completions, /v1/completions, /v1/embeddings, and /v1/ranking.
Hosted Developer Credits
1,000 free inference credits granted on NVIDIA Developer Program signup, consumed across all hosted models.
Hosted Concurrent Requests
Concurrent in-flight requests against the shared hosted endpoint.
SSE Keepalive
Streaming connections idle longer than ~60s may be closed by the gateway.
Per-request Output Cap
Per-request output token cap on the hosted endpoint (model-dependent; some models support higher).
Hosted Input Size
Maximum input tokens on the hosted endpoint (model-dependent — long-context Llama 3.1/3.3 and Nemotron variants accept up to ~128K).
Self-hosted Container Concurrency
Self-hosted NIM containers accept as many concurrent requests as the GPU and TensorRT-LLM batching configuration allow. Operators tune via NIM_MAX_BATCH_SIZE and similar env vars.

Work with this as data

Every rate limit here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for rate limits

4 MCP tools reach this
  • find_rate_limitsBrowse and filter every rate limit in the catalog.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools

Call it yourself

curl for this page
This rate limit
curl "https://apis.io/api/v1/rate-limits/nvidia-nim-rate-limits"
All rate limits
curl "https://apis.io/api/v1/rate-limits?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no email required.

A second provider on the same verified email joins the account you already have.