Cartesia Ai Rate Limits
Cartesia enforces concurrency-based limits rather than classic requests-per-minute quotas. Each plan caps the number of simultaneous TTS requests, simultaneous STT requests, and active agent slots; a single TTS WebSocket connection can multiplex "dozens" of concurrent generation contexts within its plan's TTS concurrency ceiling. Usage above the plan's monthly credit allowance is billed or blocked depending on account configuration.
Cartesia Ai Rate Limits is the machine-readable rate-limit profile for Cartesia on the APIs.io network, conforming to the API Commons Rate Limits specification.
It captures 15 rate-limit definitions, measuring connections, agents, contexts, frames, and seconds.
The profile also includes 3 backoff/retry policies defined and response codes documented for throttled.
Tagged areas include AI, Voice AI, Text to Speech, Speech to Text, and WebSocket.
Limits
Policies
Sources
- https://cartesia.ai/pricing
- https://docs.cartesia.ai/api-reference/tts/websocket
- https://docs.cartesia.ai/api-reference/stt/websocket
Work with this as data
Every rate limit here is available over the APIs.io API and to AI agents over MCP.
MCP server
One button, every client — Claude, Cursor, VS Code and the rest.
https://apis.io/mcp
Tools for rate limits
4 MCP tools reach this
find_rate_limitsBrowse and filter every rate limit in the catalog.apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.resolveTurn a domain, URL or GitHub org into the provider it belongs to.find_cohortsEvery scored population of providers in the catalog.
Call it yourself
curl for this page
curl "https://apis.io/api/v1/rate-limits/cartesia-ai-rate-limits"
curl "https://apis.io/api/v1/rate-limits?limit=25"
Discovery needs no key. Ratings and market analysis are Pro.