Cumulus Inference Gateway
OpenAI-compatible HTTP inference gateway. One client works against every upstream provider; per-workflow routing rules pick the model, provider, and infrastructure, with a layered exact/prefix/semantic cache, request-level observability, and the Ion engine serving open-weight models and fine-tunes.
Documentation
Documentation
https://docs.cumuluslabs.io/
GettingStarted
https://docs.cumuluslabs.io/getting-started/
APIReference
https://docs.cumuluslabs.io/inference/overview/
Authentication
https://raw.githubusercontent.com/api-evangelist/cumulus-labs/refs/heads/main/authentication/cumulus-labs-authentication.yml