Tabby Ml Rate Limits
Tabby does not publish fixed numeric API rate limits. Because the server is self-hosted, throughput is bounded by your own hardware - primarily the GPU (or CPU) serving the completion and chat models - rather than by a vendor-imposed per-minute request cap. Completion latency and tokens-per-second depend on the model size and device you configure. On hosted Team and Enterprise plans, capacity is governed by seat count and the managed infrastructure rather than a documented request-rate limit.
Tabby Ml Rate Limits is the machine-readable rate-limit profile for Tabby on the APIs.io network, conforming to the API Commons Rate Limits specification.
It captures 4 rate-limit definitions, measuring requests, documents, and users.
The profile also includes 3 backoff/retry policies defined and response codes documented for notImplemented.
Tagged areas include AI Coding Assistant, Code Completion, Open Source, Rate Limiting, and Quotas.