Vellum Rate Limits
Scaffolded rate limit definitions for the Vellum AI API surface. Captures per-tier quotas, burst behavior, response signaling, and recovery semantics. Defaults are scaffold values to be replaced with published provider limits.
Vellum Rate Limits is the machine-readable rate-limit profile for Vellum AI on the APIs.io network, conforming to the API Commons Rate Limits specification.
It captures 2 rate-limit definitions, across the free and pro tiers, measuring requests_per_minute.
The profile also includes response codes documented for throttled, quotaExceeded, and serviceUnavailable.
Tagged areas include LLM Platform, Prompt Engineering, Workflows, Evaluations, and LLM Ops.
Limits
Work with this as data
Every rate limit here is available over the APIs.io API and to AI agents over MCP.