vLLM website screenshot

vLLM

vLLM is a high-throughput, memory-efficient open-source inference and serving engine for LLMs. It provides an OpenAI-compatible REST server (vllm serve) plus a Python API. vLLM is Apache 2.0 and run on your own GPU infrastructure; there is no hosted vLLM SaaS from the project itself.

vLLM publishes 6 APIs on the APIs.io network, including Audio API, Chat API, Completions API, and 3 more. Tagged areas include LLM, Inference, Open-Source, GPU, and OpenAI-Compatible.

vLLM’s developer surface includes authentication, engineering blog, and 10 more developer resources.

28.4/100 thin ▬ flat Agent 20/100 agent aware saas Full breakdown ↓
scored 2026-09-14 · rubric v0.22.0
1 APIs
LLMInferenceOpen-SourceGPUOpenAI-CompatibleSelf-Hosted

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-09-14 · rubric v0.22.0
Create-or-Update Ergonomics applies to this provider. This API accepts writes, so it carries 10 points of the composite. It is scored from the published contracts themselves: whether a caller can create-or-update in one call, whether the write accepts a key the caller already holds, and whether the response says which branch ran. Without that, every write needs a search-and-branch in front of it, and the first time that check is skipped a duplicate record is created. Scored against the observed mean rather than raw — a provider at the catalog average is unchanged by this facet, not penalised by it.
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. Every facet and dimension name above is a link: it opens that measurement's own page — what it means, the exact checks that feed it, how the whole catalog distributes on it, and the providers at the top of it. This rating is computed from github.com/api-evangelist/vllm: open an issue to ask a question, or submit a pull request to add artifacts. Submit an artifact on GitHub — free → Manage your own listing — the Influence plan, $499/mo →

Standards implemented 1

Interfaces this provider implements that became standards by being copied rather than ratified. Each is profiled by the API Commons, and the evidence column says how the claim was established — not that it was made.

declared — publishes a spec that declares them
2 core of 338 operations graded · api-commons/models

APIs 7

Individual APIs this provider publishes, each with its own machine-readable definition.

vLLM OpenAI-Compatible Server

OpenAI-compatible REST API exposed by `vllm serve`. Endpoints include /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/score, /v1/audio/transcriptions, /v1/audio/trans...

vLLM Audio API

OpenAI-compatible audio endpoints

vLLM Chat API

OpenAI-compatible chat completions

vLLM Completions API

OpenAI-compatible text completions

vLLM Embeddings API

OpenAI-compatible embeddings

vLLM Scoring API

vLLM-specific scoring and reranking endpoints

vLLM Tokenize API

vLLM-specific tokenize/detokenize utilities

Scroll for all 7

Open Collections 8

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

API Collection

OPEN COLLECTION

Scroll for all 8

Pricing Plans 1

Published pricing tiers and plan structures.

Vllm Plans Pricing

1 plans

PLANS

Rate Limits 1

Documented rate limits and quota policies.

Vllm Rate Limits

2 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals for API financial operations.

Vllm Finops

FINOPS

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Vllm Authentication

http · 1 scheme

SECURITY

Vllm Domain Security

TLSv1.3

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Vllm Agentic Access

12 operations · 12 acting

12 operations · 12 acting

AGENTIC

Resources

Get Started 1

Portal, sign-up, and the first successful call

Agent Surfaces 2

MCP servers, agent skills, and machine-readable catalogs

Access & Security 2

Authentication, authorization, and security posture

Operate 1

Status, limits, changes, and where to get help

Commercial 2

Pricing, plans, and the legal terms of use

Company 3

The organization behind the API

Other 1

Properties that don't map to a standard resource type

Source (apis.yml)

apis.yml Raw ↑
aid: vllm
url: https://raw.githubusercontent.com/api-evangelist/vllm/refs/heads/main/apis.yml
name: vLLM
kind: company
description: vLLM is a high-throughput, memory-efficient open-source inference and serving engine for LLMs. It provides an
  OpenAI-compatible REST server (vllm serve) plus a Python API. vLLM is Apache 2.0 and run on your own GPU infrastructure;
  there is no hosted vLLM SaaS from the project itself.
deliveryModel:
  model: saas
  open_source: false
  commercial: true
  callable_host: false
  label: Hosted service · you call their endpoint
  confidence: medium
  source:
  - openapi
  - pricing
  generated: '2026-08-28'
  method: derived
accessModel:
  pricing: unknown
  onboarding: unknown
  trial: false
  try_now: false
  public: false
  label: Unknown
  confidence: low
  try_now_blocked_by: loopback
  source:
  - plans
  - authentication
  - security
  generated: '2026-09-03'
  method: derived
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/vllm.png
tags:
- LLM
- Inference
- Open-Source
- GPU
- OpenAI-Compatible
- Self-Hosted
tags_raw:
- LLM
- Inference
- Open Source
- GPU
- OpenAI Compatible
- Self-Hosted
created: '2026-05-08'
modified: '2026-05-08'
specificationVersion: '0.23'
apis:
- aid: vllm:openai-compatible
  name: vLLM OpenAI-Compatible Server
  description: OpenAI-compatible REST API exposed by `vllm serve`. Endpoints include /v1/chat/completions, /v1/completions,
    /v1/embeddings, /v1/score, /v1/audio/transcriptions, /v1/audio/translations, /v1/realtime (WebSocket), /tokenize, /detokenize,
    and /generative_scoring. Authentication via the --api-key flag set on server start; clients can use the official OpenAI
    Python library unmodified, with vLLM-specific extensions passed via extra_body.
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Chat
  - Completions
  - Embeddings
  - Audio
  - Score
  - OpenAI-Compatible
  properties:
  - type: Documentation
    url: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  - type: GitHub
    url: https://github.com/vllm-project/vllm
  - type: OpenAICompat
    url: https://platform.openai.com/docs/api-reference
- aid: vllm:vllm-audio-api
  name: vLLM Audio API
  description: OpenAI-compatible audio endpoints
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Audio
  properties:
  - type: OpenAPI
    url: openapi/vllm-audio-api-openapi.yml
- aid: vllm:vllm-chat-api
  name: vLLM Chat API
  description: OpenAI-compatible chat completions
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Chat
  properties:
  - type: OpenAPI
    url: openapi/vllm-chat-api-openapi.yml
- aid: vllm:vllm-completions-api
  name: vLLM Completions API
  description: OpenAI-compatible text completions
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Completions
  properties:
  - type: OpenAPI
    url: openapi/vllm-completions-api-openapi.yml
- aid: vllm:vllm-embeddings-api
  name: vLLM Embeddings API
  description: OpenAI-compatible embeddings
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Embeddings
  properties:
  - type: OpenAPI
    url: openapi/vllm-embeddings-api-openapi.yml
- aid: vllm:vllm-scoring-api
  name: vLLM Scoring API
  description: vLLM-specific scoring and reranking endpoints
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Scoring
  properties:
  - type: OpenAPI
    url: openapi/vllm-scoring-api-openapi.yml
- aid: vllm:vllm-tokenize-api
  name: vLLM Tokenize API
  description: vLLM-specific tokenize/detokenize utilities
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Tokenize
  properties:
  - type: OpenAPI
    url: openapi/vllm-tokenize-api-openapi.yml
common:
- type: AgenticAccess
  url: agentic-access/vllm-agentic-access.yml
- type: DomainSecurity
  url: security/vllm-domain-security.yml
- type: Authentication
  url: authentication/vllm-authentication.yml
- type: LinkedIn
  url: https://www.linkedin.com/company/vllm-project
- type: Website
  url: https://docs.vllm.ai/
- type: DeveloperPortal
  url: https://docs.vllm.ai/
- type: OpenSource
  url: https://github.com/vllm-project/vllm
- type: Plans
  url: plans/vllm-plans-pricing.yml
- type: RateLimits
  url: rate-limits/vllm-rate-limits.yml
- type: FinOps
  url: finops/vllm-finops.yml
- type: LlmsText
  url: https://vllm.ai/llms.txt
- url: https://vllm.ai/blog/rss.xml
  type: Blog
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com

Work with this as data

Every provider here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for providers

9 MCP tools reach this
  • find_providersBrowse and filter every provider in the catalog.
  • get_provider_artifactsEvery artifact this provider publishes, grouped by type.
  • get_provider_operationsEvery operation across all of their OpenAPIs — one call instead of parsing every spec.
  • get_provider_toolsEvery MCP tool they ship, with the operation each wraps.
  • get_provider_evidenceHow each part of their score was established. Free — the basis for a claim should not sit behind it.
  • get_provider_ratingPRO — composite, band, trend and facet scores.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This provider
curl "https://apis.io/api/v1/providers/vllm"
All providers
curl "https://apis.io/api/v1/providers?limit=25"
Every operation they expose
curl "https://apis.io/api/v1/providers/vllm/operations?limit=25"
How their score was established
curl "https://apis.io/api/v1/providers/vllm/evidence"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.