Feature

LLM Inference

Hosted model inference endpoints.

14 providers 190 APIs 13 declared variants

The provider serves text, chat, or completion inference against hosted models, usually with streaming and token-based billing. Integrators care about model choice, context window, and latency.

14 API providers on the APIs.io network offer llm inference. The highest-rated are Databricks, ChatGPT, NVIDIA NIM, Hugging Face, Hyperbolic.

Providers

Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.

Exemplar 2 Complete, well-documented, and agent-ready
Strong 7 Solid coverage with minor gaps
Developing 3 Usable, with meaningful gaps to close
Thin 2 Limited public surface area

What Providers Actually Declared

This feature is a canonical term. These are the free-text strings providers wrote in their own apis.yml that map onto it.

100+ foundation models exposed through a single API contract — Llama 3.1/3.2/3.3, Mistral, Mixtral, NVIDIA Nemotron, DeepSeek-R1, Qwen 2.5, Microsoft Phi, Google Gemma, IBM Granite, and FalconAI Model ServingChat CompletionsCoinbase x402 payment integration for crypto-native chat completionsModel InferenceModel Serving for low-latency inferenceModel serving endpoints for real-time inferenceNative tensor and ML model inference at serving timeService APIs for motion planning, vision, SLAM, navigation, data management, ML model inference, sensors aggregation, world state, discovery, and generic servicesSingle `docker run` install bundling Vespa storage and embedding model inferenceSingle-endpoint /v2/guard API following the OpenAI chat completions message formatText and Chat CompletionsUniversal LLM API

Scroll for all 13