Google Vertex Ai Rate Limits
Machine-readable rate limit definitions for the Google Vertex AI API surface. Captures per-tier quotas, burst behavior, response signaling, and recovery semantics. Defaults are scaffold values to be replaced with published provider limits.
Google Vertex Ai Rate Limits is the machine-readable rate-limit profile for Google Vertex AI on the APIs.io network, conforming to the API Commons Rate Limits specification.
It captures 5 rate-limit definitions, across the free, professional, and enterprise tiers, measuring requests_per_minute and requests_per_month.
The profile also includes 4 backoff/retry policies defined and response codes documented for throttled, quotaExceeded, and serviceUnavailable.
Tagged areas include Artificial Intelligence, Generative AI, Google Cloud, Machine Learning, and ML Models.
Limits
Policies
Work with this as data
Every rate limit here is available over the APIs.io API and to AI agents over MCP.