Together AI · Rate Limits

Together Ai Rate Limits

Together AI enforces per-account rate limits on serverless inference that vary by model and account tier (Build / Scale / Enterprise as account spend/credit grows). Limits include requests-per-minute (RPM) and tokens-per-minute (TPM) per model. Specific per-model values are not reconciled in this artifact - see the Together console for active limits on your account.

Together Ai Rate Limits is the machine-readable rate-limit profile for Together AI on the APIs.io network, conforming to the API Commons Rate Limits specification.

It captures 5 rate-limit definitions, measuring requests, tokens, and jobs.

The profile also includes 2 backoff/retry policies defined and response codes documented for throttled.

Tagged areas include AI, LLM, Inference, Open Source, and Fine-tuning.

5 Limits Throttle: 429
AILLMInferenceOpen SourceFine-tuningRate LimitingQuotasThrottling

Limits

Requests Per Minute (RPM) account
requests
see provider documentation
Per-model RPM, varies by tier and model. Pending reconciliation.
Tokens Per Minute (TPM) account
tokens
see provider documentation
Per-model TPM, varies by tier and model. Pending reconciliation.
Concurrent Fine-Tuning Jobs account
jobs
see provider documentation
Concurrency cap on parallel fine-tuning jobs.
Batch Job Size / Concurrency account
jobs
see provider documentation
Batch jobs are queued and do not consume serverless RPM/TPM directly.
Dedicated Endpoints endpoint
requests
bounded by provisioned GPU capacity
Throughput is determined by the dedicated hardware sizing.

Policies

Tiered Limits
Limits scale up automatically with account spend / credit balance and via Enterprise agreements.
Backoff Strategy
Clients should implement exponential backoff with jitter and honor any Retry-After header.

Sources

Work with this as data

Every rate limit here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for rate limits

4 MCP tools reach this
  • find_rate_limitsBrowse and filter every rate limit in the catalog.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This rate limit
curl "https://apis.io/api/v1/rate-limits/together-ai-rate-limits"
All rate limits
curl "https://apis.io/api/v1/rate-limits?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.