The provider serves text, chat, or completion inference against hosted models, usually with streaming and token-based billing. Integrators care about model choice, context window, and latency.
14 API providers on the APIs.io network offer llm inference. The highest-rated are NVIDIA NIM, Azure Databricks, Viam, Hugging Face, Hyperbolic.
Providers
Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.
Exemplar1Complete, well-documented, and agent-ready
This feature is a canonical term. These are the free-text strings
providers wrote in their own apis.yml that map onto it.
100+ foundation models exposed through a single API contract — Llama 3.1/3.2/3.3, Mistral, Mixtral, NVIDIA Nemotron, DeepSeek-R1, Qwen 2.5, Microsoft Phi, Google Gemma, IBM Granite, and FalconAI Model ServingChat CompletionsCoinbase x402 payment integration for crypto-native chat completionsModel InferenceModel Serving for low-latency inferenceModel serving endpoints for real-time inferenceNative tensor and ML model inference at serving timeService APIs for motion planning, vision, SLAM, navigation, data management, ML model inference, sensors aggregation, world state, discovery, and generic servicesSingle `docker run` install bundling Vespa storage and embedding model inferenceSingle-endpoint /v2/guard API following the OpenAI chat completions message formatText and Chat CompletionsUniversal LLM API
Every feature here is available over the APIs.io API and to AI agents over MCP. Features is not yet its own endpoint on the v1 API. Reach it through catalog search and the tag graph, or the MCP server.
MCP server
One button, every client — Claude, Cursor, VS Code and the rest.
https://apis.io/mcp
Tools for features
3 MCP tools reach this
apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
resolveTurn a domain, URL or GitHub org into the provider it belongs to.
find_cohortsEvery scored population of providers in the catalog.