Cumulus Labs website screenshot

Cumulus Labs

Cumulus Labs is a Y Combinator (W26) company building a unified inference platform for production AI. Cumulus consolidates the eight subsystems teams normally assemble from separate vendors — an OpenAI-compatible gateway, a per-workflow router, a layered prompt/KV cache, request-level observability, continuous shadow evaluation, one-click LoRA fine-tuning, custom open-weight hosting, and the proprietary Ion inference engine running on NVIDIA Grace and Blackwell GPUs — behind a single OpenAI-compatible API at api.cumuluslabs.io/v1. Ion's custom attention kernels deliver 30-50% more throughput than stock vLLM and SGLang, and the gateway is a drop-in replacement for the OpenAI, Anthropic, LangChain, LlamaIndex, and Vercel AI SDKs — change one line of configuration and keep your existing code.

Cumulus Labs publishes 1 API on the APIs.io network. Tagged areas include Company, Inference, LLM, AI Infrastructure, and GPU.

Cumulus Labs’ developer surface includes documentation, getting-started guide, API reference, engineering blog, support, authentication, and 12 more developer resources.

25.8/100 emerging ▬ flat Agent 3/100 human only saas Full breakdown ↓
scored 2026-09-08 · rubric v0.20.0
1 APIs
CompanyInferenceLLMAI InfrastructureGPUMachine-LearningModel ServingFine-TuningAPI GatewayY Combinator

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-09-08 · rubric v0.20.0
Create-or-Update Ergonomics could not be measured. We hold no machine-readable contract for this provider to read, so there is nothing to measure a write surface against. Excluded rather than scored zero: never-measured and measured-empty are different facts. Publishing an OpenAPI is what makes this facet — and several others — scorable at all.
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. Every facet and dimension name above is a link: it opens that measurement's own page — what it means, the exact checks that feed it, how the whole catalog distributes on it, and the providers at the top of it. This rating is computed from github.com/api-evangelist/cumulus-labs: open an issue to ask a question, or submit a pull request to add artifacts. Submit an artifact on GitHub — free → Manage your own listing — the Influence plan, $499/mo →

APIs 1

Individual APIs this provider publishes, each with its own machine-readable definition.

Cumulus Inference Gateway

OpenAI-compatible HTTP inference gateway. One client works against every upstream provider; per-workflow routing rules pick the model, provider, and infrastructure, with a layer...

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Cumulus Labs Authentication

http · 1 scheme

SECURITY

Cumulus Labs Domain Security

TLSv1.3 · DNSSEC · DMARC

SECURITY

Resources

Get Started 2

Portal, sign-up, and the first successful call

Documentation 2

Reference material describing how the API behaves

Agent Surfaces 1

MCP servers, agent skills, and machine-readable catalogs

Design & Contract 3

Pagination, idempotency, versioning, errors, and events

Build 1

SDKs, sample code, and the tooling you integrate with

Access & Security 2

Authentication, authorization, and security posture

Operate 1

Status, limits, changes, and where to get help

Commercial 2

Pricing, plans, and the legal terms of use

Company 4

The organization behind the API

Source (apis.yml)

apis.yml Raw ↑
aid: cumulus-labs
name: Cumulus Labs
description: Cumulus Labs is a Y Combinator (W26) company building a unified inference platform for production AI. Cumulus
  consolidates the eight subsystems teams normally assemble from separate vendors — an OpenAI-compatible gateway, a per-workflow
  router, a layered prompt/KV cache, request-level observability, continuous shadow evaluation, one-click LoRA fine-tuning,
  custom open-weight hosting, and the proprietary Ion inference engine running on NVIDIA Grace and Blackwell GPUs — behind
  a single OpenAI-compatible API at api.cumuluslabs.io/v1. Ion's custom attention kernels deliver 30-50% more throughput than
  stock vLLM and SGLang, and the gateway is a drop-in replacement for the OpenAI, Anthropic, LangChain, LlamaIndex, and Vercel
  AI SDKs — change one line of configuration and keep your existing code.
url: https://raw.githubusercontent.com/api-evangelist/cumulus-labs/refs/heads/main/apis.yml
x-type: company
x-source: vc-portfolio
x-backed-by:
- y-combinator
- nvidia-inception
x-tier: profiled
x-tier-reason: enrichment-pass
deliveryModel:
  model: saas
  open_source: false
  commercial: true
  callable_host: false
  label: Hosted service · you call their endpoint
  confidence: medium
  source:
  - pricing
  generated: '2026-08-28'
  method: derived
accessModel:
  pricing: unknown
  onboarding: unknown
  trial: false
  try_now: false
  public: false
  label: Unknown
  confidence: low
  source:
  - authentication
  - security
  generated: '2026-09-03'
  method: derived
specificationVersion: '0.23'
created: '2026-07-17'
modified: '2026-07-18'
tags:
- Company
- Inference
- LLM
- AI Infrastructure
- GPU
- Machine-Learning
- Model Serving
- Fine-Tuning
- API Gateway
- Y Combinator
tags_raw:
- Company
- Inference
- LLM
- AI Infrastructure
- GPU
- Machine Learning
- Model Serving
- Fine-Tuning
- API Gateway
- Y Combinator
image: https://cumuluslabs.io/cumulus-logo/Cumulus-White.svg
apis:
- name: Cumulus Inference Gateway
  description: OpenAI-compatible HTTP inference gateway. One client works against every upstream provider; per-workflow routing
    rules pick the model, provider, and infrastructure, with a layered exact/prefix/semantic cache, request-level observability,
    and the Ion engine serving open-weight models and fine-tunes.
  humanURL: https://docs.cumuluslabs.io/
  baseURL: https://api.cumuluslabs.io/v1
  tags:
  - Inference
  - LLM
  - OpenAI-Compatible
  - Gateway
  - GPU
  properties:
  - type: Documentation
    url: https://docs.cumuluslabs.io/
  - type: GettingStarted
    url: https://docs.cumuluslabs.io/getting-started/
  - type: APIReference
    url: https://docs.cumuluslabs.io/inference/overview/
  - type: Authentication
    url: authentication/cumulus-labs-authentication.yml
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com
- FN: APIs.json
  email: info@apis.io
common:
- type: DomainSecurity
  url: security/cumulus-labs-domain-security.yml
- type: Website
  url: https://cumuluslabs.io/
- type: DeveloperPortal
  url: https://docs.cumuluslabs.io/
- type: Documentation
  url: https://docs.cumuluslabs.io/
- type: GettingStarted
  url: https://docs.cumuluslabs.io/getting-started/
- type: APIReference
  url: https://docs.cumuluslabs.io/inference/overview/
- type: Blog
  url: https://cumulus.blog/
- type: GitHubOrganization
  url: https://github.com/cumulus-compute-labs
- type: Support
  url: mailto:founders@cumuluslabs.io
- type: TermsOfService
  url: https://cumuluslabs.io/terms-of-service.pdf
- type: PrivacyPolicy
  url: https://cumuluslabs.io/privacy-policy.pdf
- type: LinkedIn
  url: https://www.linkedin.com/company/cumuluscomputelabs/
- type: Twitter
  url: https://x.com/cumuluslabsio
- type: Authentication
  url: authentication/cumulus-labs-authentication.yml
- type: Conventions
  url: conventions/cumulus-labs-conventions.yml
- type: Conformance
  url: conformance/cumulus-labs-conformance.yml
- type: Lifecycle
  url: lifecycle/cumulus-labs-lifecycle.yml
- type: LLMsTxt
  url: llms/cumulus-labs-llms.txt
x-enrichment:
  date: '2026-07-19'
  status: backfilled
  pass: local-v1
  note: backfilled from .gitignore signal + verified work evidence

Work with this as data

Every provider here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for providers

9 MCP tools reach this
  • find_providersBrowse and filter every provider in the catalog.
  • get_provider_artifactsEvery artifact this provider publishes, grouped by type.
  • get_provider_operationsEvery operation across all of their OpenAPIs — one call instead of parsing every spec.
  • get_provider_toolsEvery MCP tool they ship, with the operation each wraps.
  • get_provider_evidenceHow each part of their score was established. Free — the basis for a claim should not sit behind it.
  • get_provider_ratingPRO — composite, band, trend and facet scores.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This provider
curl "https://apis.io/api/v1/providers/cumulus-labs"
All providers
curl "https://apis.io/api/v1/providers?limit=25"
Every operation they expose
curl "https://apis.io/api/v1/providers/cumulus-labs/operations?limit=25"
How their score was established
curl "https://apis.io/api/v1/providers/cumulus-labs/evidence"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.