Feature
LLM Inference
Hosted model inference endpoints.
14 providers
190 APIs
13 declared variants
The provider serves text, chat, or completion inference against hosted models, usually with streaming and token-based billing. Integrators care about model choice, context window, and latency.
14 API providers on the APIs.io network offer llm inference. The highest-rated are Databricks, ChatGPT, NVIDIA NIM, Hugging Face, Hyperbolic.
Providers
Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.
Exemplar 2 Complete, well-documented, and agent-ready
Strong 7 Solid coverage with minor gaps
NVIDIA NIM
NVIDIA NIM (NVIDIA Inference Microservices) is a catalog of GPU-accelerated, containerized AI inference mic...
Hugging Face
The AI community building the future with open-source machine learning models, datasets, and applications.
Hyperbolic
Hyperbolic is an open-access AI cloud and decentralized GPU marketplace serving 200,000+ builders with affo...
Viam
Viam is a robotics and edge AI platform founded in 2020 by Eliot Horowitz (MongoDB co-founder and former CT...
Lakera
Lakera is an AI-native security platform that protects GenAI applications, agents, and workforces from prom...
Vespa
Vespa is an open-source AI search engine, big-data serving engine, and vector database originally developed...
Azure Databricks
Azure Databricks is an Apache Spark-based analytics platform optimized for Microsoft Azure. It provides a c...
Developing 3 Usable, with meaningful gaps to close
Kong
Kong is the AI Connectivity Company. Its platform spans Kong Gateway (the open-source API gateway built on ...
Advanced Micro Devices
Advanced Micro Devices (AMD) is a global semiconductor company that develops high-performance computing, gr...
AIMLAPI
AIMLAPI is a unified AI model API gateway providing access to 400+ state-of-the-art AI models from OpenAI, ...
What Providers Actually Declared
This feature is a canonical term. These are the free-text strings
providers wrote in their own apis.yml that map onto it.
100+ foundation models exposed through a single API contract — Llama 3.1/3.2/3.3, Mistral, Mixtral, NVIDIA Nemotron, DeepSeek-R1, Qwen 2.5, Microsoft Phi, Google Gemma, IBM Granite, and FalconAI Model ServingChat CompletionsCoinbase x402 payment integration for crypto-native chat completionsModel InferenceModel Serving for low-latency inferenceModel serving endpoints for real-time inferenceNative tensor and ML model inference at serving timeService APIs for motion planning, vision, SLAM, navigation, data management, ML model inference, sensors aggregation, world state, discovery, and generic servicesSingle `docker run` install bundling Vespa storage and embedding model inferenceSingle-endpoint /v2/guard API following the OpenAI chat completions message formatText and Chat CompletionsUniversal LLM API
Scroll for all 13