Feature

LLM Inference

Hosted model inference endpoints.

14 providers 63 APIs 13 declared variants

The provider serves text, chat, or completion inference against hosted models, usually with streaming and token-based billing. Integrators care about model choice, context window, and latency.

14 API providers on the APIs.io network offer llm inference. The highest-rated are NVIDIA NIM, Azure Databricks, Viam, Hugging Face, Hyperbolic.

Providers

Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.

Exemplar 1 Complete, well-documented, and agent-ready
Strong 11 Solid coverage with minor gaps
Developing 2 Usable, with meaningful gaps to close

What Providers Actually Declared

This feature is a canonical term. These are the free-text strings providers wrote in their own apis.yml that map onto it.

100+ foundation models exposed through a single API contract — Llama 3.1/3.2/3.3, Mistral, Mixtral, NVIDIA Nemotron, DeepSeek-R1, Qwen 2.5, Microsoft Phi, Google Gemma, IBM Granite, and FalconAI Model ServingChat CompletionsCoinbase x402 payment integration for crypto-native chat completionsModel InferenceModel Serving for low-latency inferenceModel serving endpoints for real-time inferenceNative tensor and ML model inference at serving timeService APIs for motion planning, vision, SLAM, navigation, data management, ML model inference, sensors aggregation, world state, discovery, and generic servicesSingle `docker run` install bundling Vespa storage and embedding model inferenceSingle-endpoint /v2/guard API following the OpenAI chat completions message formatText and Chat CompletionsUniversal LLM API

Scroll for all 13

Work with this as data

Every feature here is available over the APIs.io API and to AI agents over MCP. Features is not yet its own endpoint on the v1 API. Reach it through catalog search and the tag graph, or the MCP server.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for features

3 MCP tools reach this
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
Search the catalog
curl "https://apis.io/api/v1/search?q=llm-inference&limit=10"
Everything under a tag
curl "https://apis.io/api/v1/tags/llm-inference"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no email required.

A second provider on the same verified email joins the account you already have.