Cumulus Inference Gateway
OpenAI-compatible HTTP inference gateway. One client works against every upstream provider; per-workflow routing rules pick the model, provider, and infrastructure, with a layered exact/prefix/semantic cache, request-level observability, and the Ion engine serving open-weight models and fine-tunes.