Glhf Chat Rate Limits
glhf serves an OpenAI-compatible API backed by an auto-scaling GPU scheduler. Public materials note that the service initially launched without rate limiting and that limits have since been applied, but specific per-account or per-model request and token limits are not publicly documented. As an OpenAI-compatible surface, throttling is expected to be returned as HTTP 429. Specific values are not reconciled in this artifact.
Glhf Chat Rate Limits is the machine-readable rate-limit profile for glhf on the APIs.io network, conforming to the API Commons Rate Limits specification.
It captures 3 rate-limit definitions, measuring requests and tokens.
The profile also includes 1 backoff/retry policy defined and response codes documented for throttled.
Tagged areas include AI, LLM, Inference, Open Source Models, and Hugging Face.
Limits
Policies
Sources
Work with this as data
Every rate limit here is available over the APIs.io API and to AI agents over MCP.