vLLM website screenshot

vLLM

vLLM is a high-throughput, memory-efficient open-source inference and serving engine for LLMs. It provides an OpenAI-compatible REST server (vllm serve) plus a Python API. vLLM is Apache 2.0 and run on your own GPU infrastructure; there is no hosted vLLM SaaS from the project itself.

vLLM publishes 6 APIs on the APIs.io network, including Audio API, Chat API, Completions API, and 3 more. Tagged areas include LLM, Inference, Open-Source, GPU, and OpenAI-Compatible.

vLLM’s developer surface includes authentication, engineering blog, and 10 more developer resources.

27.0/100 thin ▬ flat Agent 20/100 agent aware Full breakdown ↓
scored 2026-08-25 · rubric v0.14.0
AccessSelf serve
7 APIs
LLMInferenceOpen-SourceGPUOpenAI-CompatibleSelf-Hosted

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-08-25 · rubric v0.14.0
Composite quality — 27.0/100 · thin
Contract Quality 11.7 / 25
Developer Ergonomics 4.8 / 20
Access Clarity 2.6 / 20
Operational Transparency 0.7 / 13
Contract Governance 0.0 / 12
Discoverability 7.2 / 10
Agent readiness — 20/100 · agent aware
Machine-Readable Contract 18 / 18
Agentic Access Contract 10 / 10
Documented Reversibility 0 / 6
MCP Server 0 / 12
Machine-Readable Auth 10 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 0 / 7
Rate-Limit Signaling 7 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3
A2A Agent Card 0 / 8
Dry-Run / Simulate Mode 0 / 4
Delegated User Identity 0 / 6
Protected Resource Metadata 0 / 5
Registration Without a Human 0 / 6
Agentic Commerce Surface 0 / 5
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. This rating is computed from github.com/api-evangelist/vllm: open an issue to ask a question, or submit a pull request to add artifacts. Want it done for you? Prioritized profiling — $2,500 →

APIs 7

Individual APIs this provider publishes, each with its own machine-readable definition.

vLLM OpenAI-Compatible Server

OpenAI-compatible REST API exposed by `vllm serve`. Endpoints include /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/score, /v1/audio/transcriptions, /v1/audio/trans...

vLLM Audio API

OpenAI-compatible audio endpoints

vLLM Chat API

OpenAI-compatible chat completions

vLLM Completions API

OpenAI-compatible text completions

vLLM Embeddings API

OpenAI-compatible embeddings

vLLM Scoring API

vLLM-specific scoring and reranking endpoints

vLLM Tokenize API

vLLM-specific tokenize/detokenize utilities

Scroll for all 7

Open Collections 8

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

API Collection

OPEN COLLECTION

Scroll for all 8

Pricing Plans 1

Published pricing tiers and plan structures.

Vllm Plans Pricing

1 plans

PLANS

Rate Limits 1

Documented rate limits and quota policies.

Vllm Rate Limits

2 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals for API financial operations.

Vllm Finops

FINOPS

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Vllm Authentication

http · 1 scheme

SECURITY

Vllm Domain Security

TLSv1.3

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Vllm Agentic Access

12 operations · 12 acting

12 operations · 12 acting

AGENTIC

Resources

Get Started 1

Portal, sign-up, and the first successful call

Agent Surfaces 2

MCP servers, agent skills, and machine-readable catalogs

Access & Security 2

Authentication, authorization, and security posture

Operate 1

Status, limits, changes, and where to get help

Commercial 2

Pricing, plans, and the legal terms of use

Company 3

The organization behind the API

Other 1

Properties that don't map to a standard resource type

Source (apis.yml)

apis.yml Raw ↑
aid: vllm
url: https://raw.githubusercontent.com/api-evangelist/vllm/refs/heads/main/apis.yml
name: vLLM
kind: company
description: vLLM is a high-throughput, memory-efficient open-source inference and serving engine for LLMs. It provides an
  OpenAI-compatible REST server (vllm serve) plus a Python API. vLLM is Apache 2.0 and run on your own GPU infrastructure;
  there is no hosted vLLM SaaS from the project itself.
accessModel:
  pricing: unknown
  onboarding: self-serve
  trial: false
  try_now: false
  public: false
  label: Self-serve signup
  confidence: high
  source:
  - plans
  - authentication
  generated: '2026-07-22'
  method: derived
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/vllm.png
tags:
- LLM
- Inference
- Open-Source
- GPU
- OpenAI-Compatible
- Self-Hosted
tags_raw:
- LLM
- Inference
- Open Source
- GPU
- OpenAI Compatible
- Self-Hosted
created: '2026-05-08'
modified: '2026-05-08'
specificationVersion: '0.23'
apis:
- aid: vllm:openai-compatible
  name: vLLM OpenAI-Compatible Server
  description: OpenAI-compatible REST API exposed by `vllm serve`. Endpoints include /v1/chat/completions, /v1/completions,
    /v1/embeddings, /v1/score, /v1/audio/transcriptions, /v1/audio/translations, /v1/realtime (WebSocket), /tokenize, /detokenize,
    and /generative_scoring. Authentication via the --api-key flag set on server start; clients can use the official OpenAI
    Python library unmodified, with vLLM-specific extensions passed via extra_body.
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Chat
  - Completions
  - Embeddings
  - Audio
  - Score
  - OpenAI-Compatible
  properties:
  - type: Documentation
    url: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  - type: GitHub
    url: https://github.com/vllm-project/vllm
  - type: OpenAICompat
    url: https://platform.openai.com/docs/api-reference
- aid: vllm:vllm-audio-api
  name: vLLM Audio API
  description: OpenAI-compatible audio endpoints
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Audio
  properties:
  - type: OpenAPI
    url: openapi/vllm-audio-api-openapi.yml
- aid: vllm:vllm-chat-api
  name: vLLM Chat API
  description: OpenAI-compatible chat completions
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Chat
  properties:
  - type: OpenAPI
    url: openapi/vllm-chat-api-openapi.yml
- aid: vllm:vllm-completions-api
  name: vLLM Completions API
  description: OpenAI-compatible text completions
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Completions
  properties:
  - type: OpenAPI
    url: openapi/vllm-completions-api-openapi.yml
- aid: vllm:vllm-embeddings-api
  name: vLLM Embeddings API
  description: OpenAI-compatible embeddings
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Embeddings
  properties:
  - type: OpenAPI
    url: openapi/vllm-embeddings-api-openapi.yml
- aid: vllm:vllm-scoring-api
  name: vLLM Scoring API
  description: vLLM-specific scoring and reranking endpoints
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Scoring
  properties:
  - type: OpenAPI
    url: openapi/vllm-scoring-api-openapi.yml
- aid: vllm:vllm-tokenize-api
  name: vLLM Tokenize API
  description: vLLM-specific tokenize/detokenize utilities
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Tokenize
  properties:
  - type: OpenAPI
    url: openapi/vllm-tokenize-api-openapi.yml
common:
- type: AgenticAccess
  url: agentic-access/vllm-agentic-access.yml
- type: DomainSecurity
  url: security/vllm-domain-security.yml
- type: Authentication
  url: authentication/vllm-authentication.yml
- type: LinkedIn
  url: https://www.linkedin.com/company/vllm-project
- type: Website
  url: https://docs.vllm.ai/
- type: DeveloperPortal
  url: https://docs.vllm.ai/
- type: OpenSource
  url: https://github.com/vllm-project/vllm
- type: Plans
  url: plans/vllm-plans-pricing.yml
- type: RateLimits
  url: rate-limits/vllm-rate-limits.yml
- type: FinOps
  url: finops/vllm-finops.yml
- type: LlmsText
  url: https://vllm.ai/llms.txt
- url: https://vllm.ai/blog/rss.xml
  type: Blog
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com

Work with this as data

Every provider here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for providers

9 MCP tools reach this
  • find_providersBrowse and filter every provider in the catalog.
  • get_provider_artifactsEvery artifact this provider publishes, grouped by type.
  • get_provider_operationsEvery operation across all of their OpenAPIs — one call instead of parsing every spec.
  • get_provider_toolsEvery MCP tool they ship, with the operation each wraps.
  • get_provider_evidenceHow each part of their score was established. Free — the basis for a claim should not sit behind it.
  • get_provider_ratingPRO — composite, band, trend and facet scores.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools

Call it yourself

curl for this page
This provider
curl "https://apis.io/api/v1/providers/vllm"
All providers
curl "https://apis.io/api/v1/providers?limit=25"
Every operation they expose
curl "https://apis.io/api/v1/providers/vllm/operations?limit=25"
How their score was established
curl "https://apis.io/api/v1/providers/vllm/evidence"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no email required.

A second provider on the same verified email joins the account you already have.