established · 1.5 capabilitymarket actiondomain

Inference

Score 29.7 / 100 149 providers 80 APIs Search apis.io →
Variants seen in the corpus: Inferenceinference

Providers using this tag (149)

Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.

Exemplar 2 Complete, well-documented, and agent-ready
Strong 18 Solid coverage with minor gaps
Developing 43 Usable, with meaningful gaps to close
InferenceCompanyArtificial IntelligenceMachine-LearningLLM1 API53.9SambaNova SystemsCompanyArtificial IntelligenceMachine-LearningLLM12 APIs53.8SimplismartCompanyArtificial IntelligenceMachine-Learning8 APIs53.3RunPodArtificial IntelligenceCloudComputeGPU10 APIs52.2ParasailArtificial IntelligenceGPULarge Language Models7 APIs51.6Fastino LabsCompanyArtificial IntelligenceMachine-LearningSmall Language Models4 APIs51.50G LabsArtificial IntelligenceAI InferenceLLMGPU Compute9 APIs51.2LambdaArtificial IntelligenceCloudClusterCompute13 APIs50.9OpenRelayCompanyGPUArtificial Intelligence19 APIs50.8Sail ResearchCompanyArtificial IntelligenceLLM5 APIs50.2FriendliAICompanyInfrastructureArtificial IntelligenceMachine-Learning33 APIs49.9OpenworkCompanyAI AgentsOpen-SourceDesktop37 APIs49.7Pruna AICompanyArtificial IntelligenceMachine-LearningImage-Generation3 APIs48.1SailCompanyArtificial IntelligenceMachine-LearningLLM5 APIs48.1Aleph AlphaCompanyArtificial IntelligenceMachine-LearningLarge Language Models37 APIs47.9RouterPlexLLMArtificial IntelligenceAI Gateway5 APIs47.2Recursal AI, Inc.CompanyArtificial IntelligenceMachine-LearningLLM3 APIs47.0FlexAICompanyAi MlArtificial IntelligenceMachine-Learning7 APIs46.8Hugging Face TransformersArtificial IntelligenceComputer-VisionDeep LearningMachine-Learning27 APIs45.3Sahara AICompanyCryptoArtificial IntelligenceMachine-Learning2 APIs45.2TensorWaveCompanyArtificial IntelligenceMachine-LearningCloud Computing5 APIs45.1FuriosaAIArtificial IntelligenceMachine-LearningSemiconductors3 APIs44.6CoreWeaveArtificial IntelligenceCloudGPUHPC6 APIs44.2Novita AIArtificial IntelligenceLLMGPU2 APIs44.2BentoMLMachine-LearningModel ServingArtificial Intelligence62 APIs44.1AI21 LabsArtificial IntelligenceFoundation ModelsLLMJamba11 APIs43.9Adaptive MLCompanyAi MlLLMFine-Tuning9 APIs43.5.txtCompanyArtificial IntelligenceLLMStructured Outputs6 APIs42.4BeamServerlessGPUPython3 APIs42.1CerebrasAI InferenceLarge Language ModelsWafer ScaleHardware4 APIs42.0Featherless AIArtificial IntelligenceLLMServerless4 APIs42.0ChutesArtificial IntelligenceLLMServerless4 APIs41.7NscaleArtificial IntelligenceGPUServerless5 APIs41.6SUTRA (Two AI)Artificial IntelligenceLLMMultilingual2 APIs41.4Arcee AICompanyArtificial IntelligenceMachine-LearningLarge Language Models29 APIs41.2RunwareCompanyArtificial IntelligenceMachine-Learning1 API41.0OxenCompanyData Version ControlMachine-LearningArtificial Intelligence19 APIs40.8CentMLArtificial IntelligenceLLMServerless5 APIs40.7Amazon NovaFoundation ModelsGenerative AIImage-GenerationMachine-Learning3 APIs40.4OllamaArtificial IntelligenceLarge Language ModelsModels14 APIs40.4Akash NetworkCloud ComputingDecentralizedBlockchainKubernetes41 APIs39.9PredibaseArtificial IntelligenceLLMFine-Tuning7 APIs39.9TargonArtificial IntelligenceLLMDecentralized5 APIs39.4
Thin 55 Limited public surface area
SambaNovaAI InferenceLarge Language ModelsDataflowsHardware5 APIs38.9InferlessArtificial IntelligenceML InferenceServerless GPUModel Deployment2 APIs38.6SeldonMLOpsMachine-LearningModel Serving31 APIs38.6Azure OpenAI ServiceArtificial IntelligenceLLMGenerative AIAzure8 APIs38.5TensorFlowArtificial IntelligenceDeep LearningJavaScriptMachine-Learning7 APIs38.2CerebriumArtificial IntelligenceGPUServerless3 APIs38.0glhfArtificial IntelligenceLLMOpen Source Models2 APIs38.0DataCrunchGPU CloudInfrastructureCompute11 APIs37.9OvershootCompanyArtificial IntelligenceComputer-VisionVideo8 APIs37.9Secton APIArtificial IntelligenceLLMChat Completions2 APIs37.6Scalable Inference ServingArtificial IntelligenceCNCFDeployment9 APIs37.5GcoreEdge CloudCDNStreamingEdge AI8 APIs37.2Qubrid AIArtificial IntelligenceCloud ComputingGPU13 APIs37.2Weights and BiasesMLOpsExperiment TrackingLLM ObservabilityModel Registry16 APIs36.7BasetenArtificial IntelligenceMLDeployment3 APIs35.6LaminiArtificial IntelligenceLLMFine-TuningMemory Tuning5 APIs35.3Together AIArtificial IntelligenceLLMFoundation Models28 APIs35.1GroqArtificial IntelligenceLLMLPU9 APIs34.3TaalasCompanyArtificial IntelligenceAI InferenceSemiconductors3 APIs34.3ModularCompanyArtificial IntelligenceMachine-Learning1 API34.0CrusoeAI CloudGPUCompute48 APIs33.7poolsideCompanyEnterpriseArtificial IntelligenceMachine-Learning1 API33.7EigenLayerRestakingAVSEthereumData Availability7 APIs33.5overshoot.aiCompanyArtificial IntelligenceComputer-VisionVideo8 APIs33.4Fireworks AIArtificial IntelligenceLLMMulti-Modal10 APIs33.1Triton Inference ServerArtificial IntelligenceDeep LearningMachine-Learning12 APIs33.1Thinking MachinesCompanyArtificial IntelligenceMachine-LearningFine-Tuning1 API33.0Moonshot AIArtificial IntelligenceLLMLong Context6 APIs32.3Inworld AIArtificial IntelligenceVoiceCharactersGames10 APIs31.9AI GatewayAI GatewayLLM RouterLLM ProxyModel Routing38 APIs31.8WomboCompanyArtificial IntelligenceMachine-Learning4 APIs31.5GoodfireCompanyArtificial IntelligenceMachine-LearningInterpretability1 API31.4SAP Business Technology PlatformSAPCloud PlatformIntegrationArtificial Intelligence7 APIs30.9FluidstackArtificial IntelligenceGPUCloudCompute7 APIs30.5MiniMaxArtificial IntelligenceLLMMulti-Modal7 APIs30.3ModelbitCompanyArtificial IntelligenceMachine-LearningMLOps1 API29.7MangoBoostCompanyArtificial IntelligenceMachine-LearningInfrastructure3 APIs29.6LM StudioCompanyArtificial IntelligenceLocal LLMMachine-Learning4 APIs29.5KServeKubernetesMachine-LearningMLOps4 APIs29.4QwenArtificial IntelligenceLLMOpen-Source4 APIs29.301.AIArtificial IntelligenceLLMYiOpen-Source2 APIs29.2Zhipu AIArtificial IntelligenceLLMGLM3 APIs29.0DeepInfraArtificial IntelligenceLLMServerless7 APIs28.9HyperbolicArtificial IntelligenceLLMGPU7 APIs28.7ByteDance DoubaoArtificial IntelligenceLLMByteDance6 APIs27.9LatticaCompanyPrivacyFully Homomorphic EncryptionEncryption1 API27.9SiliconFlowArtificial IntelligenceLLMOpen-Source11 APIs27.8OpenPipeArtificial IntelligenceLLMFine-TuningDistillation10 APIs27.7Lightning AICompanyAi MlMachine-LearningGPU Cloud2 APIs27.2StriveworksCompanyArtificial IntelligenceMachine-LearningMLOps1 API27.0vLLMLLMOpen-SourceGPU7 APIs27.0AnyscaleArtificial IntelligenceDistributed ComputingRayML Platform10 APIs26.9AsterlabCompanyArtificial IntelligenceMachine-LearningLLM1 API26.5TensorZeroCompanyAi MlLLMLLMOps26.5KAITOArtificial IntelligenceGPUKubernetes5 APIs26.3
Emerging 18 Early or largely undocumented
Minimal 13 Almost no public developer surface

APIs with this tag (77)

Ranked by the provider's API Evangelist rating — the Kin Score is scored per provider, not per API, so every API of a provider shares its band. How the rating works →

Exemplar 2 Complete, well-documented, and agent-ready
Strong 12 Solid coverage with minor gaps
Developing 18 Usable, with meaningful gaps to close
Inference.net APIOpenAI-compatible inference API for open-source, frontier, and custom language models — chat completions, b...53.9RunPod ServerlessRunPod Serverless provides pay-as-you-go inference endpoints with autoscaling workers, queue-based and load...52.2Fastino Labs inference APIPioneer-native inference endpoint (encoder NER/classification/extraction and decoder text generation).51.50G Labs Inference APIThe Inference API from 0G Labs — 11 operation(s) for inference.51.2Lambda Inference APILambda Inference API is an OpenAI-compatible REST gateway at https://api.lambda.ai/v1 that serves hosted op...50.9OpenRelay Inference APIOpenAI-compatible chat completions API (POST /v1/chat/completions) and, for supporting models, an Anthropic...50.8Openwork Inference APIThe Inference API from Openwork — 1 operation(s) for inference.49.7Hugging Face Inference APIServerless inference API for running predictions against thousands of models hosted on the Hugging Face Hub...45.3Sahara AI Inference APIOpenAI-compatible chat completion inference.45.2CoreWeave Inference APIThe CoreWeave Inference API manages Deployments, Gateways, and Capacity Claims for serverless and dedicated...44.2BentoML Service REST APIAuto-generated REST API endpoints produced when BentoML services are deployed. Each decorated service metho...44.1AI21 Batch APIAsynchronous batch processing for large volumes of Jamba completions. Submit a batch job, poll for status, ...43.9Cerebras Inference APIThe Cerebras Inference API exposes ultra-low-latency inference for open-weight large language models includ...42.0Runware Inference APISingle task-based endpoint for image, video, audio, 3D, and text inference across 400K+ models, reachable o...41.0Amazon Nova Inference APISynchronous and streaming inference operations.40.4Ollama Cloud APIOllama Cloud provides cloud-hosted inference for large language models, giving access to larger models and ...40.4AkashML Inference APIOpenAI-compatible AI inference API providing access to open-source LLMs hosted on decentralized compute acr...39.9Predibase Inference APIThe Inference API from Predibase — 4 operation(s) for inference.39.9
Thin 36 Limited public surface area
SambaCloud APIThe SambaCloud API exposes OpenAI-compatible chat completions over SambaNova's RDU-accelerated infrastructu...38.9Inferless Inference APIThe Inference API from Inferless — 1 operation(s) for inference.38.6Seldon Inference APIThe Seldon Inference API provides REST and gRPC endpoints for serving machine learning model predictions at...38.6Azure OpenAI Inference REST APIData-plane REST API for running inference against deployed Azure OpenAI models, including chat completions,...38.5TensorFlow Inference APIModel inference operations including classify, regress, and predict38.2Cerebrium Inference APIThe Inference API from Cerebrium — 1 operation(s) for inference.38.0BentoML REST APIBentoML is an open-source unified inference platform for deploying and scaling AI models. It auto-generates...37.5NVIDIA Triton Inference Server HTTP APINVIDIA Triton Inference Server is an open-source inference serving software that implements the KServe Open...37.5Ray Serve REST APIRay Serve is a scalable model serving library built on Ray, designed for building online inference APIs. Su...37.5Scalable Inference Serving Inference APIModel inference request endpoints37.5vLLM OpenAI-Compatible APIvLLM is a high-throughput and memory-efficient inference engine for LLMs, implementing PagedAttention for e...37.5Gcore Inference APIEverywhere Inference edge AI model deployments.37.2W&B Serverless Inference (CoreWeave)OpenAI-compatible inference API for hosted open-source foundation models, running on CoreWeave GPU infrastr...36.7Lamini Inference APIText completion and streaming generation endpoints.35.3Taalas APITaalas-native REST interface for running inference against the HC1 hardcore-model silicon. Three operations...34.3MAX Inference REST APIOpenAI-compatible inference API served by MAX. Exposes /v1/chat/completions, /v1/completions, /v1/embedding...34.0Poolside APIOpenAI-compatible inference API for poolside's Laguna agentic-coding models. Send chat-completion and model...33.7EigenAIEigenAI exposes deterministic, verifiable LLM inference over an OpenAI-compatible API surface, with executi...33.5Triton GRPC APIHigh-performance gRPC API for model inference with support for streaming and binary tensor data.33.1Triton Inference Server Inference APIModel inference requests33.1Tinker APIManaged training/fine-tuning and sampling API for open-weight language models, consumed through the officia...33.0Inworld LLM Router APIRouting layer over 220+ LLM models, billed at provider cost via Inworld's unified API.31.9AnyscaleAnyscale is the production-scale AI platform built on Ray by the creators of Ray, supporting LLM inference ...31.8NVIDIA NIMNVIDIA NIM is a set of inference microservices for streamlined AI model deployment, prebuilt and optimized ...31.8Together AITogether AI is a full-stack AI Native Cloud for inference, fine-tuning, and GPU clusters powered by researc...31.8Ember APIHosted mechanistic-interpretability API. OpenAI-compatible chat-completions sampling plus feature discovery...31.4SAP AI Core APIREST API for SAP AI Core, enabling training, deployment, and inference of AI and machine learning models wi...30.9Modelbit Deployment REST APIEvery Modelbit deployment is exposed as a versioned REST inference endpoint. POST an inference request (sin...29.7Mango LLMBoost Inference Server APILLMBoost is MangoBoost's enterprise LLM inference server. It serves the OpenAI REST API on /v1 so an existi...29.6KServe Inference APIKServe's standardized model inference protocol for serving predictions across multiple ML frameworks on Kub...29.4Lattica Platform APIThe LatticaAI platform API for deploying and operating encrypted workloads. An RPC-style HTTPS surface root...27.9Lightning AI Model APIsHosted inference API that serves frontier and open models behind a Lightning API key, billed per token with...27.2Chariot Platform APIPer-tenant REST API for the Chariot AI operations platform, organized as versioned microservice paths under...27.0Anyscale Services APIDeploys and manages production Ray Serve applications as long-running, autoscaling, multi-version services ...26.9Aster Inference APIOpenAI-compatible inference API from Aster serving open-weight models (gpt-oss-120b, gpt-oss-120b-fast, GLM...26.5KAITO RAGEngine APIRAGEngine exposes endpoints for managing retrieval-augmented generation services with embedded vector datab...26.3
Emerging 8 Early or largely undocumented
Minimal 1 Almost no public developer surface

Score breakdown

Frequency
64.2
log-scaled weighted occurrences
Breadth
4.1
spread across providers
Quality lift
43.8
mean composite of providers using it
Cohesion
25.7
strength of nearest seed neighbor

Related tags

OpenAI-Compatible 42 co-occurrences GPU 45 co-occurrences LLM 80 co-occurrences Models 51 co-occurrences Completions 25 co-occurrences Fine-Tuning 23 co-occurrences Chat 42 co-occurrences Embeddings 26 co-occurrences

Where this tag comes from

Provider tag116
Api tag79
Openapi tag20
Openapi op tag60

Cohort brief

Auto-generated

The 149 providers in the APIs.io catalog tagged Inference, scored on the Kin Score. Every figure below is computed from the catalog — nothing here is written.

Providers
149
all scored
Mean Kin Score
36.3
+15.6 vs catalog 20.7
Mean Agent Readiness
20.2
+10.5 vs catalog 9.7
Spread
5–71.2
median 37.2 · σ 15.6
How the 149 split by band
Exemplar 2Strong 18Developing 43Thin 55Emerging 18Minimal 13
Facet averages, against the whole catalog
FacetThis cohortCatalogDifferenceScored
Contract Quality 40.1 17.5 +22.6 149
Developer Ergonomics 40.1 17.9 +22.2 149
Operational Transparency 23.4 11.3 +12.1 149
Discoverability 72.1 60.4 +11.7 149
Access Clarity 32.4 22.3 +10.1 149
Contract Governance 11.1 6.4 +4.7 149

A facet is averaged over the members that carry it, not over the whole cohort — the “Scored” column is that count. Averaging an absent facet as zero would score our own coverage gaps as the providers’ posture.

What this cohort publishes
ArtifactThis cohortCatalogDifference
MCP server (any) 30% 16% +14
MCP server (first-party) 13% 7% +6
Agent Skills 0% 0% 0
OAuth scopes 5% 9% -4
Security 96% 88% +8
Arazzo workflows 6% 2% +4
Governance rules 26% 13% +13

mcp_pct counts any mcp/ artifact including ones API Evangelist derived from the provider OpenAPI; mcp_first_party_pct counts only servers the provider publishes. Prefer the latter.

Top by Kin Score
  1. 1 NVIDIA NIM 71.2
  2. 2 fal 69.7
  3. 3 Autoderm – AI Dermatology API 66.2
  4. 4 Modal 65.1
  5. 5 Amazon SageMaker 65
  6. 6 Hyperbolic 64.7
  7. 7 Cerebras Systems 63.8
  8. 8 Swisscom 63.8
  9. 9 Viam 61.6
  10. 10 Compresr 61.4
Top by Agent Readiness
  1. 1 Fastino Labs 54.2
  2. 2 Roboflow 54
  3. 3 Hugging Face Transformers 53.9
  4. 4 Openwork 52.6
  5. 5 Swisscom 48
  6. 6 H Company 43.3
  7. 7 Dedalus Labs 42.8
  8. 8 OpenRelay 40.6
  9. 9 Crusoe 40.3
  10. 10 Sail 40
All 149 members of this roster resolve to a scored provider in the catalog.Generated from the catalog build of 25 August 2026, across 26891 providers, using the same computation served by the APIs.io cohort API. Briefs are published for rosters of 5 or more scored providers; below that a distribution is not meaningful.

Work with this as data

Every tag here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for tags

7 MCP tools reach this
  • find_tagsBrowse and filter every tag in the catalog.
  • get_cohortThis tag as a scored cohort — every provider carrying it, with scores.
  • cohort_statsPRO — the distribution across this tag: mean, median, band split, adoption rates.
  • cohort_rankingsPRO — the leaderboard, on composite AND agent-readiness axes.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools

Call it yourself

curl for this page
This tag
curl "https://apis.io/api/v1/tags/inference"
All tags
curl "https://apis.io/api/v1/tags?limit=25"
As a scored cohort
curl "https://apis.io/api/v1/cohorts/tag/inference"
The distribution (Pro)
curl "https://apis.io/api/v1/cohorts/tag/inference/stats" \
  -H "X-API-Key: $APIS_IO_KEY"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no email required.

A second provider on the same verified email joins the account you already have.