Gradient Labs Rate Limits
Gradient Labs does not publish specific numeric rate limits for its HTTP API on its public site or in its open-source SDKs. The API documents that endpoints are idempotent and requests can be safely retried, which implies clients should implement retry-with-backoff behavior. Any account-level throttling and quotas are expected to be governed by the enterprise agreement. Webhook deliveries carry an incrementing sequence_number per conversation so consumers can order events and safely deduplicate on retry.
Gradient Labs Rate Limits is the machine-readable rate-limit profile for Gradient Labs on the APIs.io network, conforming to the API Commons Rate Limits specification.
It captures 2 rate-limit definitions, measuring requests and events.
The profile also includes 2 backoff/retry policies defined and response codes documented for throttled.
Tagged areas include AI, Customer Support, AI Agent, Financial Services, and Rate Limiting.
Limits
Policies
Sources
Work with this as data
Every rate limit here is available over the APIs.io API and to AI agents over MCP.