Cerebras Systems
Cerebras Systems builds the Wafer-Scale Engine (WSE) — the largest computer chip ever made — and the CS-3 systems built around it, delivering AI training and inference at speeds far beyond conventional GPUs. Cerebras Inference is the company's cloud API: an OpenAI-compatible REST interface at api.cerebras.ai that serves open-weight frontier models (OpenAI GPT-OSS, Gemma 4, Z.ai GLM 4.7) at industry-leading tokens-per-second. Developers authenticate with a bearer API key and call chat/completions, completions, and models endpoints, with streaming, tool calling, structured outputs, vision, prompt caching, and batch inference. Custom model weights can be deployed on dedicated endpoints. Official Python and Node.js SDKs, a Cloud Console with a playground, and coding-tool integrations (VS Code, Cline, Kilo Code, OpenCode) round out the developer surface.
Cerebras Systems publishes 5 APIs on the APIs.io network, including Chat API, Completions API, Models API, and 2 more. Tagged areas include Company, Ai Infrastructure, Artificial Intelligence, Machine Learning, and Inference.
Cerebras Systems’ developer surface includes documentation, API reference, getting-started guide, support, engineering blog, pricing, signup flow, and 33 more developer resources.
Kin Score
APIs 5
Individual APIs this provider publishes, each with its own machine-readable definition.
Cerebras Systems Chat API
The Chat API from Cerebras Systems — 1 operation(s) for chat.
Cerebras Systems Completions API
The Completions API from Cerebras Systems — 1 operation(s) for completions.
Cerebras Systems Models API
The Models API from Cerebras Systems — 2 operation(s) for models.
Cerebras Systems Public Models API
The Public Models API from Cerebras Systems — 2 operation(s) for public models.
Cerebras Systems Tcp Warming API
The Tcp Warming API from Cerebras Systems — 1 operation(s) for tcp warming.
Postman Collections 5
Ready-to-run Postman collections for exercising this provider's APIs.
Cerebras Inference Chat API
POSTMANOpen Collections 6
Open, tool-agnostic API collections (OpenAPI-derived and Bruno).
API Collection
OPEN COLLECTIONCerebras Inference Chat API
OPEN COLLECTIONCerebras Inference Chat Completions API
OPEN COLLECTIONCerebras Inference Chat Models API
OPEN COLLECTIONCerebras Inference Chat Public Models API
OPEN COLLECTIONCerebras Inference Chat Tcp Warming API
OPEN COLLECTIONMCP Servers 1
Model Context Protocol servers that expose these APIs to AI agents.
cerebras-systems-mcp.yml
MCP SERVERPricing Plans 1
Published pricing tiers and plan structures.
Rate Limits 1
Documented rate limits and quota policies.
Cerebras Systems Rate Limits
RATE LIMITSSecurity Posture 3
Authentication, domain security, vulnerability disclosure, and trust-center signals.
Agentic Access 1
Recommended x-agentic-access execution contracts for AI agents.
Resources
Get Started 4
Portal, sign-up, and the first successful call
Documentation 2
Reference material describing how the API behaves
Agent Surfaces 4
MCP servers, agent skills, and machine-readable catalogs
Design & Contract 5
Pagination, idempotency, versioning, errors, and events
Build 6
SDKs, sample code, and the tooling you integrate with
Access & Security 6
Authentication, authorization, and security posture
Operate 6
Status, limits, changes, and where to get help
Commercial 4
Pricing, plans, and the legal terms of use
Company 2
The organization behind the API
Other 1
Properties that don't map to a standard resource type