Cumulus Inference Gateway

OpenAI-compatible HTTP inference gateway. One client works against every upstream provider; per-workflow routing rules pick the model, provider, and infrastructure, with a layered exact/prefix/semantic cache, request-level observability, and the Ion engine serving open-weight models and fine-tunes.

API entry from apis.yml

apis.yml Raw ↑
name: Cumulus Inference Gateway
description: OpenAI-compatible HTTP inference gateway. One client works against every upstream provider;
  per-workflow routing rules pick the model, provider, and infrastructure, with a layered exact/prefix/semantic
  cache, request-level observability, and the Ion engine serving open-weight models and fine-tunes.
humanURL: https://docs.cumuluslabs.io/
baseURL: https://api.cumuluslabs.io/v1
tags:
- Inference
- LLM
- OpenAI-Compatible
- Gateway
- GPU
properties:
- type: Documentation
  url: https://docs.cumuluslabs.io/
- type: GettingStarted
  url: https://docs.cumuluslabs.io/getting-started/
- type: APIReference
  url: https://docs.cumuluslabs.io/inference/overview/
- type: Authentication
  url: authentication/cumulus-labs-authentication.yml