Feature
LLM Inference
Hosted model inference endpoints.
14 providers
386 APIs
13 declared variants
The provider serves text, chat, or completion inference against hosted models, usually with streaming and token-based billing. Integrators care about model choice, context window, and latency.
14 API providers on the APIs.io network offer llm inference. The highest-rated are NVIDIA NIM, ChatGPT, Databricks, Hugging Face, Hyperbolic.
Providers
Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.
Exemplar 4 Complete, well-documented, and agent-ready
NVIDIA NIM
NVIDIA NIM (NVIDIA Inference Microservices) is a catalog of GPU-accelerated, containerized AI inference mic...
ChatGPT
OpenAI's ChatGPT API for conversational AI and language model interactions.
Databricks
Collection of Databricks REST APIs for managing workspaces, clusters, jobs, and data operations.
Hugging Face
The AI community building the future with open-source machine learning models, datasets, and applications.
Strong 8 Solid coverage with minor gaps
Hyperbolic
Hyperbolic is an open-access AI cloud and decentralized GPU marketplace serving 200,000+ builders with affo...
Viam
Viam is a robotics and edge AI platform founded in 2020 by Eliot Horowitz (MongoDB co-founder and former CT...
Lakera
Lakera is an AI-native security platform that protects GenAI applications, agents, and workforces from prom...
Vespa
Vespa is an open-source AI search engine, big-data serving engine, and vector database originally developed...
Kong
Kong is the AI Connectivity Company. Its platform spans Kong Gateway (the open-source API gateway built on ...
Advanced Micro Devices
Advanced Micro Devices (AMD) is a global semiconductor company that develops high-performance computing, gr...
Azure Databricks
Azure Databricks is an Apache Spark-based analytics platform optimized for Microsoft Azure. It provides a c...
AIMLAPI
AIMLAPI is a unified AI model API gateway providing access to 400+ state-of-the-art AI models from OpenAI, ...
What Providers Actually Declared
This feature is a canonical term. These are the free-text strings
providers wrote in their own apis.yml that map onto it.
100+ foundation models exposed through a single API contract — Llama 3.1/3.2/3.3, Mistral, Mixtral, NVIDIA Nemotron, DeepSeek-R1, Qwen 2.5, Microsoft Phi, Google Gemma, IBM Granite, and FalconAI Model ServingChat CompletionsCoinbase x402 payment integration for crypto-native chat completionsModel InferenceModel Serving for low-latency inferenceModel serving endpoints for real-time inferenceNative tensor and ML model inference at serving timeService APIs for motion planning, vision, SLAM, navigation, data management, ML model inference, sensors aggregation, world state, discovery, and generic servicesSingle `docker run` install bundling Vespa storage and embedding model inferenceSingle-endpoint /v2/guard API following the OpenAI chat completions message formatText and Chat CompletionsUniversal LLM API
Scroll for all 13