vLLM website screenshot

vLLM

vLLM is a high-throughput, memory-efficient open-source inference and serving engine for LLMs. It provides an OpenAI-compatible REST server (vllm serve) plus a Python API. vLLM is Apache 2.0 and run on your own GPU infrastructure; there is no hosted vLLM SaaS from the project itself.

vLLM publishes 6 APIs on the APIs.io network, including Audio API, Chat API, Completions API, and 3 more. Tagged areas include LLM, Inference, Open Source, GPU, and OpenAI Compatible.

vLLM’s developer surface includes authentication, engineering blog, and 10 more developer resources.

32.7/100 thin ▬ flat Agent 31/100 agent aware Full breakdown ↓
scored 2026-08-05 · rubric v0.9.1
AccessSelf serve
7 APIs
LLMInferenceOpen SourceGPUOpenAI CompatibleSelf-Hosted

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-08-05 · rubric v0.9.1
Composite quality — 32.7/100 · thin
Contract Quality 13.4 / 25
Developer Ergonomics 4.3 / 20
Commercial Clarity 5.8 / 20
Operational Transparency 2.7 / 13
Governance 0.0 / 12
Discoverability 6.5 / 10
Agent readiness — 31/100 · agent aware
Machine-Readable Contract 18 / 18
Agentic Access Contract 10 / 10
MCP Server 0 / 12
Machine-Readable Auth 10 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 0 / 7
Rate-Limit Signaling 7 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3
A2A Agent Card 0 / 8
Dry-Run / Simulate Mode 0 / 4
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. This rating is computed from github.com/api-evangelist/vllm: open an issue to ask a question, or submit a pull request to add artifacts. Want it done for you? Prioritized profiling — $2,500 →

APIs 7

Individual APIs this provider publishes, each with its own machine-readable definition.

vLLM OpenAI-Compatible Server

OpenAI-compatible REST API exposed by `vllm serve`. Endpoints include /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/score, /v1/audio/transcriptions, /v1/audio/trans...

vLLM Audio API

OpenAI-compatible audio endpoints

vLLM Chat API

OpenAI-compatible chat completions

vLLM Completions API

OpenAI-compatible text completions

vLLM Embeddings API

OpenAI-compatible embeddings

vLLM Scoring API

vLLM-specific scoring and reranking endpoints

vLLM Tokenize API

vLLM-specific tokenize/detokenize utilities

Scroll for all 7

Open Collections 1

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

Pricing Plans 1

Published pricing tiers and plan structures.

Vllm Plans Pricing

1 plans

PLANS

Rate Limits 1

Documented rate limits and quota policies.

Vllm Rate Limits

2 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals for API financial operations.

Vllm Finops

FINOPS

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Vllm Authentication

http · 1 scheme

SECURITY

Vllm Domain Security

TLSv1.3

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Vllm Agentic Access

12 operations · 12 acting

12 operations · 12 acting

AGENTIC

Resources

Get Started 1

Portal, sign-up, and the first successful call

Agent Surfaces 2

MCP servers, agent skills, and machine-readable catalogs

Access & Security 2

Authentication, authorization, and security posture

Operate 1

Status, limits, changes, and where to get help

Commercial 2

Pricing, plans, and the legal terms of use

Company 3

The organization behind the API

Other 1

Properties that don't map to a standard resource type

Source (apis.yml)

apis.yml Raw ↑
aid: vllm
url: https://raw.githubusercontent.com/api-evangelist/vllm/refs/heads/main/apis.yml
name: vLLM
kind: company
description: vLLM is a high-throughput, memory-efficient open-source inference and serving engine for LLMs. It provides an
  OpenAI-compatible REST server (vllm serve) plus a Python API. vLLM is Apache 2.0 and run on your own GPU infrastructure;
  there is no hosted vLLM SaaS from the project itself.
accessModel:
  pricing: unknown
  onboarding: self-serve
  trial: false
  try_now: false
  public: false
  label: Self-serve signup
  confidence: high
  source:
  - plans
  - authentication
  generated: '2026-07-22'
  method: derived
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/vllm.png
tags:
- LLM
- Inference
- Open Source
- GPU
- OpenAI Compatible
- Self-Hosted
created: '2026-05-08'
modified: '2026-05-08'
specificationVersion: '0.19'
apis:
- aid: vllm:openai-compatible
  name: vLLM OpenAI-Compatible Server
  description: OpenAI-compatible REST API exposed by `vllm serve`. Endpoints include /v1/chat/completions, /v1/completions,
    /v1/embeddings, /v1/score, /v1/audio/transcriptions, /v1/audio/translations, /v1/realtime (WebSocket), /tokenize, /detokenize,
    and /generative_scoring. Authentication via the --api-key flag set on server start; clients can use the official OpenAI
    Python library unmodified, with vLLM-specific extensions passed via extra_body.
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Chat
  - Completions
  - Embeddings
  - Audio
  - Score
  - OpenAI-Compatible
  properties:
  - type: Documentation
    url: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  - type: GitHub
    url: https://github.com/vllm-project/vllm
  - type: OpenAICompat
    url: https://platform.openai.com/docs/api-reference
- aid: vllm:vllm-audio-api
  name: vLLM Audio API
  description: OpenAI-compatible audio endpoints
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Audio
  properties:
  - type: OpenAPI
    url: openapi/vllm-audio-api-openapi.yml
- aid: vllm:vllm-chat-api
  name: vLLM Chat API
  description: OpenAI-compatible chat completions
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Chat
  properties:
  - type: OpenAPI
    url: openapi/vllm-chat-api-openapi.yml
- aid: vllm:vllm-completions-api
  name: vLLM Completions API
  description: OpenAI-compatible text completions
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Completions
  properties:
  - type: OpenAPI
    url: openapi/vllm-completions-api-openapi.yml
- aid: vllm:vllm-embeddings-api
  name: vLLM Embeddings API
  description: OpenAI-compatible embeddings
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Embeddings
  properties:
  - type: OpenAPI
    url: openapi/vllm-embeddings-api-openapi.yml
- aid: vllm:vllm-scoring-api
  name: vLLM Scoring API
  description: vLLM-specific scoring and reranking endpoints
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Scoring
  properties:
  - type: OpenAPI
    url: openapi/vllm-scoring-api-openapi.yml
- aid: vllm:vllm-tokenize-api
  name: vLLM Tokenize API
  description: vLLM-specific tokenize/detokenize utilities
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Tokenize
  properties:
  - type: OpenAPI
    url: openapi/vllm-tokenize-api-openapi.yml
common:
- type: AgenticAccess
  url: agentic-access/vllm-agentic-access.yml
- type: DomainSecurity
  url: security/vllm-domain-security.yml
- type: Authentication
  url: authentication/vllm-authentication.yml
- type: LinkedIn
  url: https://www.linkedin.com/company/vllm-project
- type: Website
  url: https://docs.vllm.ai/
- type: DeveloperPortal
  url: https://docs.vllm.ai/
- type: OpenSource
  url: https://github.com/vllm-project/vllm
- type: Plans
  url: plans/vllm-plans-pricing.yml
- type: RateLimits
  url: rate-limits/vllm-rate-limits.yml
- type: FinOps
  url: finops/vllm-finops.yml
- type: LlmsText
  url: https://vllm.ai/llms.txt
- url: https://vllm.ai/blog/rss.xml
  type: Blog
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com