established · 0.4 capabilitymarket actiondomain

Inference

Score 29.5 / 100 157 providers +37 via APIs 107 APIs Search apis.io →
Workflows using this tag
TensorFlow Serving Preflight and Predict · TensorFlow TensorFlow Serving Preflight and Regress · TensorFlow TensorFlow Serving Preflight and Classify · TensorFlow TensorFlow Serving Compare a Candidate Version Against the Default · TensorFlow TensorFlow Serving Pinned Reproducible Example Scoring · TensorFlow TensorFlow Serving Route Inference by Version Label · TensorFlow TensorFlow Serving Gate a Rollout on Version Readiness · TensorFlow Viam Machine Health Check · Viam
Business capabilities this tag reaches
Information & Data Management 15
Variants seen in the corpus: AI InferenceLLM InferenceInference APIInference

Providers using this tag (157)

Ranked by API Evangelist rating — Exemplar and Strong are expanded by default.

Exemplar 6 Complete, well-documented, and agent-ready
Strong 12 Solid coverage with minor gaps
Developing 41 Usable, with meaningful gaps to close
FriendliAICompanyInfrastructureArtificial IntelligenceMachine Learning1 API53.8ReplicateArtificial IntelligenceMachine LearningImage GenerationLanguage Models1 API53.1RightNow AICompanyArtificial IntelligenceMachine Learning1 API53.1SambaNova SystemsCompanyArtificial IntelligenceMachine LearningLLM2 APIs52.7Dedalus LabsCompanyArtificial IntelligenceAgentsMCP2 APIs52.6LambdaArtificial IntelligenceCloudClusterCompute1 API51.7TypeSafe AIArtificial IntelligenceMachine LearningClassificationContent Moderation2 APIs51.6TARXANCompanyAI AgentsAgent RuntimeLocal-First AI3 APIs50.3RunPodArtificial IntelligenceCloudComputeGPU1 API49.8ParasailArtificial IntelligenceGPULLM3 APIs49.50G LabsArtificial IntelligenceLLMGPU Compute2 APIs48.7TensorWaveCompanyArtificial IntelligenceMachine LearningCloud Computing1 API48.0Fastino LabsCompanyArtificial IntelligenceMachine LearningSmall Language Models1 API47.9Multiverse ComputingArtificial IntelligenceMachine LearningModel Compression2 APIs47.9Akash NetworkCloud ComputingDecentralizedBlockchainKubernetes3 APIs47.5Sail ResearchCompanyArtificial IntelligenceLLM1 API47.5ModelRushArtificial IntelligenceLLMMulti-Modal1 API47.1Standard ComputeLLM APIFlat RateSubscriptionAI Agents1 API46.9Novita AIArtificial IntelligenceLLMGPU1 API46.5Moore ThreadsCompanyGPUArtificial IntelligenceMachine Learning4 APIs46.3Pruna AICompanyArtificial IntelligenceMachine LearningImage Generation1 API45.4Aleph AlphaCompanyArtificial IntelligenceMachine LearningLLM6 APIs45.3Superb AIArtificial IntelligenceMachine LearningComputer VisionData Labeling1 API45.2SailCompanyArtificial IntelligenceMachine LearningLLM1 API44.9FlexAICompanyAi MlArtificial IntelligenceMachine Learning1 API44.8CoreWeaveArtificial IntelligenceCloudGPUHPC1 API44.7brick.blueAI AgentsAgent MarketplaceAgent DiscoveryTask Exchange2 APIs44.4RouterPlexLLMArtificial IntelligenceAI Gateway1 API43.5RunwareCompanyArtificial IntelligenceMachine Learning1 API43.5FuriosaAIArtificial IntelligenceMachine LearningSemiconductors2 APIs43.3QuantCDNCDNEdgeStatic HostingJAMstack1 API42.9Sahara AICompanyCryptoArtificial IntelligenceMachine Learning1 API42.4PaleBlueDot.AIArtificial IntelligenceMachine LearningLLM2 APIs42.3Qubrid AIArtificial IntelligenceCloud ComputingGPU4 APIs42.2BerrerAgentsAgentic CommerceA2AMCP1 API41.8AI21 LabsArtificial IntelligenceFoundation ModelsLLMJamba1 API40.6BentoMLMachine LearningModel ServingArtificial Intelligence1 API40.4Arcee AICompanyArtificial IntelligenceMachine LearningLLM1 API40.1Preferred NetworksCompanyArtificial IntelligenceMachine LearningLLM1 API40.0Adaptive MLCompanyAi MlLLMFine-Tuning1 API39.6Scalable Inference ServingArtificial IntelligenceCNCFDeployment1 API39.5
Thin 51 Limited public surface area
BeamServerlessGPUPython1 API38.9ChutesArtificial IntelligenceLLMServerless1 API38.6LocalAIArtificial IntelligenceMachine LearningLLM2 APIs38.4SUTRA (Two AI)Artificial IntelligenceLLMMultilingual1 API38.0Secton APIArtificial IntelligenceLLMChat Completions1 API37.7SeldonMLOpsMachine LearningModel Serving3 APIs37.6Triton Inference ServerArtificial IntelligenceDeep LearningMachine Learning2 APIs37.6Featherless AIArtificial IntelligenceLLMServerless2 APIs37.5.txtCompanyArtificial IntelligenceLLMStructured Outputs1 API37.2NscaleArtificial IntelligenceGPUServerless1 API37.2CerebriumArtificial IntelligenceGPUServerless1 API37.1CentMLArtificial IntelligenceLLMServerless1 API37.0OxenCompanyData Version ControlMachine LearningArtificial Intelligence2 APIs36.7ModularCompanyArtificial IntelligenceMachine Learning1 API36.5PredibaseArtificial IntelligenceLLMFine-Tuning1 API36.2TargonArtificial IntelligenceLLMDecentralized1 API36.2InferlessArtificial IntelligenceML InferenceServerless GPUModel Deployment1 API36.1DataCrunchGPU CloudInfrastructureCompute1 API35.9io.netArtificial IntelligenceGPUDecentralized ComputeDePIN8 APIs35.8OvershootCompanyArtificial IntelligenceComputer VisionVideo2 APIs35.1glhfArtificial IntelligenceLLMOpen Source Models1 API34.8Moonshot AIArtificial IntelligenceLLMLong Context1 API34.6Together AIArtificial IntelligenceLLMFoundation Models3 APIs34.5WeaveAPI - OpenAI-compatible AI API GatewayArtificial IntelligenceLLMAPI Gateway2 APIs34.0HailoArtificial IntelligenceMachine LearningSemiconductorsEdge Computing2 APIs33.3GroqArtificial IntelligenceLLMLPU1 API33.0GoodfireCompanyArtificial IntelligenceMachine LearningInterpretability1 API32.9TaalasCompanyArtificial IntelligenceSemiconductors3 APIs32.9PositronArtificial Intelligenceinference-hardwareAI Accelerators3 APIs32.7BasetenArtificial IntelligenceMachine LearningDeployment2 APIs32.6LaminiArtificial IntelligenceLLMFine-TuningMemory Tuning1 API32.6DeepInfraArtificial IntelligenceLLMServerless1 API32.4Zhipu AIArtificial IntelligenceLLMGLM1 API32.301.AIArtificial IntelligenceLLMYiOpen Source1 API31.5MiniMaxArtificial IntelligenceLLMMulti-Modal1 API31.4MangoBoostCompanyArtificial IntelligenceMachine LearningInfrastructure3 APIs30.7QwenArtificial IntelligenceLLMOpen Source1 API30.6Fireworks AIArtificial IntelligenceLLMMulti-Modal5 APIs29.9SiliconFlowArtificial IntelligenceLLMOpen Source1 API29.5WomboCompanyArtificial IntelligenceMachine Learning1 API29.1ByteDance DoubaoArtificial IntelligenceLLMByteDance1 API28.8LatticaCompanyPrivacyFully Homomorphic EncryptionEncryption1 API28.6LM StudioCompanyArtificial IntelligenceLocal LLMMachine Learning1 API28.4HyperbolicArtificial IntelligenceLLMGPU1 API27.8TensorZeroCompanyAi MlLLMLLMOps27.4FluidstackArtificial IntelligenceGPUCloudCompute1 API27.0OpenPipeArtificial IntelligenceLLMFine-TuningDistillation1 API27.0AsterlabCompanyArtificial IntelligenceMachine LearningLLM1 API26.7KServeKubernetesMachine LearningMLOps1 API26.7Cumulus LabsCompanyLLMAI Infrastructure1 API26.5vLLMLLMOpen SourceGPU1 API26.2
Emerging 21 Early or largely undocumented
Minimal 25 Almost no public developer surface
Unrated 1 Not yet scored

APIs with this tag (100)

Ranked by the provider's API Evangelist rating — the Kin Score is scored per provider, not per API, so every API of a provider shares its band. How the rating works →

Exemplar 9 Complete, well-documented, and agent-ready
Strong 18 Solid coverage with minor gaps
NexGen Cloud Inference APIOpenAI-compatible chat completions endpoint for running inference against deployed models. Use this endpoin...65.9Infer by Flow7 Inference APIAuthenticated customer inference operations.65.7Paperspace Deployments APIContainer-as-a-service deployments that run user-provided images on Paperspace GPU machines with a managed ...65.1Model Serving Endpoints APICreate and manage model serving endpoints for deploying machine learning models as REST API endpoints.63.1Seekr Inference APIThe Inference API from Seekr — 12 operation(s) for inference.63.0Amazon Nova Inference APISynchronous and streaming inference against the Amazon Nova understanding and generation models through the...62.1Compresr Inference APIThe Inference API from Compresr — 1 operation(s) for inference.61.0Autoderm – AI Dermatology API Inference APIThe inference API from Autoderm – AI Dermatology API — 5 operation(s) for inference.60.5Viam Inference APICloud-hosted inference against registry-deployed models.60.2Hugging Face Inference APIServerless inference API for running predictions against thousands of models hosted on the Hugging Face Hub...60.0Swiss AI Platform Inference Endpoints APIOpenAI-compatible inference endpoints for chat, multimodal, audio, embedding, reasoning and guardrail model...59.1Openwork Inference APIThe Inference API from Openwork — 1 operation(s) for inference.58.5Crusoe Managed Inference APIOpenAI-compatible inference API from the Crusoe Intelligence Foundry. Send chat/completions and embeddings ...56.9Pinecone Inference APIModel inference55.8OpenRelay Inference APIOpenAI-compatible chat completions API (POST /v1/chat/completions) and, for supporting models, an Anthropic...55.1Amazon Bedrock Inference APIOperations for invoking models and running inference.55.0Prime Intellect Inference APIOpenAI-compatible inference API for hosted frontier and open models served at api.pinference.ai. Supports s...55.0Inference.net APIOpenAI-compatible inference API for open-source, frontier, and custom language models — chat completions, b...54.3
Developing 27 Usable, with meaningful gaps to close
H Company Holo Models APIOpenAI-compatible inference API serving the Holo3 and Holo3.1 vision-language models for computer use — cha...51.8Lambda Inference APILambda Inference API is an OpenAI-compatible REST gateway at https://api.lambda.ai/v1 that serves hosted op...51.7TARX Supercomputer APIOpenAI-compatible hosted inference API. GET /v1/models lists the single model t-supercomputer (owned_by tar...50.3TrueFoundry Model Serving APITrueFoundry's Model Serving capability enables deployment and management of LLM and embedding models using ...50.0RunPod ServerlessRunPod Serverless provides pay-as-you-go inference endpoints with autoscaling workers, queue-based and load...49.80G Labs Inference APIThe Inference API from 0G Labs — 11 operation(s) for inference.48.7Akkio Inference APIThe Inference API from Akkio — 1 operation(s) for inference.48.6Fastino Labs Inference APIPioneer-native inference endpoint (encoder NER/classification/extraction and decoder text generation).47.9AkashML Inference APIOpenAI-compatible AI inference API providing access to open-source LLMs hosted on decentralized compute acr...47.5Sonde Platform Service APIThe Sonde Platform Service API lets partners run Sonde Vocal Biomarker Health Checks from their own mobile,...46.8Azure OpenAI Inference REST APIData-plane REST API for running inference against deployed Azure OpenAI models, including chat completions,...46.3KUAE Cloud Coding Plan APIAn LLM inference endpoint operated on Moore Threads MTT S5000 GPUs and sold as a subscription for AI coding...46.3Huma Inference APIThe Inference API from Huma — 2 operation(s) for inference.45.4CoreWeave Inference APIThe CoreWeave Inference API manages Deployments, Gateways, and Capacity Claims for serverless and dedicated...44.7Runware Inference APISingle task-based endpoint for image, video, audio, 3D, and text inference across 400K+ models, reachable o...43.5Furiosa-LLM OpenAI-Compatible ServerThe HTTP server started by `furiosa-llm serve `. It hosts a single model on RNGD NPUs and exposes an Open43.3QuantCDN AI Inference APIChat inference, embeddings, and image generation services42.9Sahara AI Inference APIOpenAI-compatible chat completion inference.42.4PBD TokenRouter Inference APIOpenAI-, Anthropic- and Gemini-compatible inference gateway. One API key and one base URL route requests ac...42.3W&B Serverless Inference (CoreWeave)OpenAI-compatible inference API for hosted open-source foundation models, running on CoreWeave GPU infrastr...41.7AI21 Batch APIAsynchronous batch processing for large volumes of Jamba completions. Submit a batch job, poll for status, ...40.6BentoML Service REST APIAuto-generated REST API endpoints produced when BentoML services are deployed. Each decorated service metho...40.4BentoML REST APIBentoML is an open-source unified inference platform for deploying and scaling AI models. It auto-generates...39.5NVIDIA Triton Inference Server HTTP APINVIDIA Triton Inference Server is an open-source inference serving software that implements the KServe Open...39.5Ray Serve REST APIRay Serve is a scalable model serving library built on Ray, designed for building online inference APIs. Su...39.5Scalable Inference Serving Inference APIModel inference request endpoints39.5vLLM OpenAI-Compatible APIvLLM is a high-throughput and memory-efficient inference engine for LLMs, implementing PagedAttention for e...39.5
Thin 34 Limited public surface area
LocalAI Inference APIThe inference API from LocalAI — 7 operation(s) for inference.38.4Ollama Cloud APIOllama Cloud provides cloud-hosted inference for large language models, giving access to larger models and ...37.8Seldon Inference APIThe Seldon Inference API provides REST and gRPC endpoints for serving machine learning model predictions at...37.6Triton GRPC APIHigh-performance gRPC API for model inference with support for streaming and binary tensor data.37.6Triton Inference Server Inference APIModel inference requests37.6AnyscaleAnyscale is the production-scale AI platform built on Ray by the creators of Ray, supporting LLM inference ...37.4NVIDIA NIMNVIDIA NIM is a set of inference microservices for streamlined AI model deployment, prebuilt and optimized ...37.4Together AITogether AI is a full-stack AI Native Cloud for inference, fine-tuning, and GPU clusters powered by researc...37.4Cerebrium Inference APIThe Inference API from Cerebrium — 1 operation(s) for inference.37.1TensorFlow Inference APIModel inference operations including classify, regress, and predict36.6MAX Inference REST APIOpenAI-compatible inference API served by MAX. Exposes /v1/chat/completions, /v1/completions, /v1/embedding...36.5Tinker APIManaged training/fine-tuning and sampling API for open-weight language models, consumed through the officia...36.3Predibase Inference APIThe Inference API from Predibase — 4 operation(s) for inference.36.2Inferless Inference APIThe Inference API from Inferless — 1 operation(s) for inference.36.1IO Intelligence APIOpenAI-compatible inference API for open-source AI models hosted on io.net's decentralized GPU network. Exp...35.8WeaveAPI Anthropic-compatible Messages APIAnthropic/Claude-wire-compatible messages route served from the bare api.weaveapi.dev host. The Claude Code...34.0WeaveAPI OpenAI-compatible APIOpenAI-wire-compatible REST API for chat completions, responses-style requests and the model catalog, authe...34.0Gcore Inference APIEverywhere Inference edge AI model deployments.33.5King's e-Research AI Hub APIAn OpenAI-compatible LLM inference API operated by King's e-Research for researchers, students and staff. P...33.5HailoRTHailoRT is Hailo's production runtime library for the Hailo-8, Hailo-10 and Hailo-15 device families. It is...33.3Ember APIHosted mechanistic-interpretability API. OpenAI-compatible chat-completions sampling plus feature discovery...32.9Taalas APITaalas-native REST interface for running inference against the HC1 hardcore-model silicon. Three operations...32.9Lamini Inference APIText completion and streaming generation endpoints.32.6TUD:AI LLM APIAn OpenAI-compatible LLM inference API operated by ZIH and ScaDS.AI Dresden/Leipzig for TU Dresden staff, s...32.6Inworld LLM Router APIRouting layer over 220+ LLM models, billed at provider cost via Inworld's unified API.31.2Mango LLMBoost Inference Server APILLMBoost is MangoBoost's enterprise LLM inference server. It serves the OpenAI REST API on /v1 so an existi...30.7EigenAIEigenAI exposes deterministic, verifiable LLM inference over an OpenAI-compatible API surface, with executi...30.4Ray Serve HTTP APIHTTP interface for invoking models and applications deployed via Ray Serve. Each deployed application is ex...29.7SAP AI Core APIREST API for SAP AI Core, enabling training, deployment, and inference of AI and machine learning models wi...29.3Lattica Platform APIThe LatticaAI platform API for deploying and operating encrypted workloads. An RPC-style HTTPS surface root...28.6Modelbit Deployment REST APIEvery Modelbit deployment is exposed as a versioned REST inference endpoint. POST an inference request (sin...27.8Aster Inference APIOpenAI-compatible inference API from Aster serving open-weight models (gpt-oss-120b, gpt-oss-120b-fast, GLM...26.7KServe Inference APIKServe's standardized model inference protocol for serving predictions across multiple ML frameworks on Kub...26.7Cumulus Inference GatewayOpenAI-compatible HTTP inference gateway. One client works against every upstream provider; per-workflow ro...26.5
Emerging 10 Early or largely undocumented
Minimal 2 Almost no public developer surface

Companies reaching this through an API (37)

These companies publish an API, specification or operation carrying “Inference” but do not classify their business under it. Listed unranked and kept out of the count above, because one tagged operation is not a statement about what a company does.

AI Gateway Akkio Amazon Bedrock Amazon Nova Autoderm – AI Dermatology API Azure Databricks Azure OpenAI Service Cloudflare Coin Railz Compresr EigenLayer Elastic Elastic Stack Gcore H Company Hugging Face Transformers Huma Infer by Flow7 Inworld AI King's College London Lightning AI Modelbit Ollama Openwork Paperspace Pinecone Ray SAP Business Technology Platform Snowflake Sonde Health Swisscom TensorFlow Thinking Machines TrueFoundry TU Dresden Viam Weights and Biases

Score breakdown

Frequency
67.4
log-scaled weighted occurrences
Breadth
4.1
spread across providers
Quality lift
40.9
mean composite of providers using it
Cohesion
22.8
strength of nearest seed neighbor

Where this tag sits

157 providers carry this tag directly, and 163 counting the 1 narrower tag below — 6 more to reach.

Artificial Intelligence broader Model Serving 14 providers

Related tags

OpenAI-Compatible 48 co-occurrences LLM 101 co-occurrences GPU 48 co-occurrences Models 60 co-occurrences Completions 25 co-occurrences Chat 43 co-occurrences Embeddings 31 co-occurrences Fine-Tuning 23 co-occurrences

Where this tag comes from

Provider tag157
Api tag107
Openapi tag28
Openapi op tag76

Cohort brief

Auto-generated

The 154 providers in the APIs.io catalog tagged Inference, scored on the Kin Score. Every figure below is computed from the catalog — nothing here is written.

Providers
154
all scored
Mean Kin Score
34
+11.1 vs catalog 22.9
Mean Agent Readiness
19.8
+7.9 vs catalog 11.9
Spread
0–74.1
median 35.8 · σ 17
How the 154 split by band
Exemplar 6Strong 12Developing 41Thin 51Emerging 21Minimal 23
Facet averages, against the whole catalog
FacetThis cohortCatalogDifferenceScored
Developer Ergonomics 42.6 23.2 +19.4 154
Contract Quality 31.8 17.6 +14.2 154
Operational Transparency 22.8 13.7 +9.1 154
Access Clarity 34.8 26.5 +8.3 154
Discoverability 66.3 58.1 +8.2 154
Contract Governance 7.3 6.3 +1 154

A facet is averaged over the members that carry it, not over the whole cohort — the “Scored” column is that count. Averaging an absent facet as zero would score our own coverage gaps as the providers’ posture.

What this cohort publishes
ArtifactThis cohortCatalogDifference
MCP server (any) 23% 14% +9
MCP server (first-party) 21% 12% +9
Agent Skills 0% 0% 0
OAuth scopes 3% 11% -8
Security 97% 98% -1
Arazzo workflows 5% 2% +3
Governance rules 19% 14% +5

mcp_pct counts any mcp/ artifact including ones API Evangelist derived from the provider OpenAPI; mcp_first_party_pct counts only servers the provider publishes. Prefer the latter.

Top by Kin Score
  1. 1 Hive Civilization 74.1
  2. 2 Segmind 68.8
  3. 3 fal 68.6
  4. 4 Hugging Face 68
  5. 5 NVIDIA NIM 68
  6. 6 Amazon SageMaker 67.3
  7. 7 NexGen Cloud 65.9
  8. 8 Seekr 63
  9. 9 Cerebras Systems 59.1
  10. 10 Hyperbolic 57.7
Top by Agent Readiness
  1. 1 brick.blue 56.6
  2. 2 Roboflow 54
  3. 3 Fastino Labs 52
  4. 4 fal 50.6
  5. 5 Replicate 50.6
  6. 6 Inference 47.1
  7. 7 NexGen Cloud 44.7
  8. 8 Hive Civilization 44.3
  9. 9 Sail 43.2
  10. 10 Superb AI 42.2
All 157 members of this roster resolve to a scored provider in the catalog.Generated from the catalog build of 5 October 2026, across 29093 providers, using the same computation served by the APIs.io cohort API. Briefs are published for rosters of 5 or more scored providers; below that a distribution is not meaningful.

Work with this as data

Every tag here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for tags

7 MCP tools reach this
  • find_tagsBrowse and filter every tag in the catalog.
  • get_cohortThis tag as a scored cohort — every provider carrying it, with scores.
  • cohort_statsPRO — the distribution across this tag: mean, median, band split, adoption rates.
  • cohort_rankingsPRO — the leaderboard, on composite AND agent-readiness axes.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This tag
curl "https://apis.io/api/v1/tags/inference"
All tags
curl "https://apis.io/api/v1/tags?limit=25"
As a scored cohort
curl "https://apis.io/api/v1/cohorts/tag/inference"
The distribution (Pro)
curl "https://apis.io/api/v1/cohorts/tag/inference/stats" \
  -H "X-API-Key: $APIS_IO_KEY"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.