Feature

LLM Inference

Hosted model inference endpoints.

13 providers 62 APIs 12 declared variants

The provider serves text, chat, or completion inference against hosted models, usually with streaming and token-based billing. Integrators care about model choice, context window, and latency.

13 API providers on the APIs.io network offer llm inference. The highest-rated are Hugging Face, NVIDIA NIM, AIMLAPI, Azure Databricks, Kong.

Providers

Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.

Exemplar 2 Complete, well-documented, and agent-ready
Strong 8 Solid coverage with minor gaps
Developing 2 Usable, with meaningful gaps to close
Thin 1 Limited public surface area

What Providers Actually Declared

This feature is a canonical term. These are the free-text strings providers wrote in their own apis.yml that map onto it.

100+ foundation models exposed through a single API contract — Llama 3.1/3.2/3.3, Mistral, Mixtral, NVIDIA Nemotron, DeepSeek-R1, Qwen 2.5, Microsoft Phi, Google Gemma, IBM Granite, and FalconAI Model ServingChat CompletionsCoinbase x402 payment integration for crypto-native chat completionsModel InferenceModel Serving for low-latency inferenceModel serving endpoints for real-time inferenceNative tensor and ML model inference at serving timeService APIs for motion planning, vision, SLAM, navigation, data management, ML model inference, sensors aggregation, world state, discovery, and generic servicesSingle `docker run` install bundling Vespa storage and embedding model inferenceText and Chat CompletionsUniversal LLM API

Scroll for all 12

Work with this as data

Every feature here is available over the APIs.io API and to AI agents over MCP. Features is not yet its own endpoint on the v1 API. Reach it through catalog search and the tag graph, or the MCP server.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for features

3 MCP tools reach this
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
Search the catalog
curl "https://apis.io/api/v1/search?q=llm-inference&limit=10"
Everything under a tag
curl "https://apis.io/api/v1/tags/llm-inference"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.