vLLM website screenshot

vLLM

vLLM is a high-throughput, memory-efficient open-source inference and serving engine for LLMs. It provides an OpenAI-compatible REST server (vllm serve) plus a Python API. vLLM is Apache 2.0 and run on your own GPU infrastructure; there is no hosted vLLM SaaS from the project itself.

vLLM publishes 7 APIs on the APIs.io network, including Audio API, Completions API, Embeddings API, and 4 more. Tagged areas include LLM, Inference, Open Source, GPU, and OpenAI-Compatible.

vLLM’s developer surface includes authentication, engineering blog, and 10 more developer resources.

26.2/100 thin ▬ flat Agent 18/100 agent aware open core · Apache-2.0 Full breakdown ↓
scored 2026-10-04 · rubric v0.23.0
1 published contract
LLMInferenceOpen SourceGPUOpenAI-CompatibleSelf-Hosted

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-10-04 · rubric v0.23.0
Regulatory Posture applies to this provider. Its tags matched the Horizontal (data, software, accessibility, platform) regime, so Regulatory Posture carries 15 points of the composite. If this regime is wrong for your business, say so on your provider repo — the applicability map is public and we will correct it.
Create-or-Update Ergonomics applies to this provider. This API accepts writes, so it carries 10 points of the composite. It is scored from the published contracts themselves: whether a caller can create-or-update in one call, whether the write accepts a key the caller already holds, and whether the response says which branch ran. Without that, every write needs a search-and-branch in front of it, and the first time that check is skipped a duplicate record is created. Scored against the observed mean rather than raw — a provider at the catalog average is unchanged by this facet, not penalised by it.
The six quality facets above are damped to 75 points between them, because the conditional facet above carries the other 25. That is why each facet's contribution is shown against a damped maximum: raising a quality facet moves the composite by 75% of its nominal weight, not 100%. The full arithmetic is at apis.io/rating/.
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. Every facet and dimension name above is a link: it opens that measurement's own page — what it means, the exact checks that feed it, how the whole catalog distributes on it, and the providers at the top of it. This rating is computed from github.com/api-evangelist/vllm: open an issue to ask a question, or submit a pull request to add artifacts. Submit an artifact on GitHub — free → Manage your own listing — the Influence plan, $499/mo →

Standards implemented 1

Interfaces this provider implements that became standards by being copied rather than ratified. Each is profiled by the API Commons, and the evidence column says how the claim was established — not that it was made.

declared — publishes a spec that declares them
2 core of 338 operations graded · api-commons/models

APIs 7

Individual APIs this provider publishes, each with its own machine-readable definition.

vLLM OpenAI-Compatible Server

OpenAI-compatible REST API exposed by `vllm serve`. Endpoints include /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/score, /v1/audio/transcriptions, /v1/audio/trans...

vLLM Audio API

OpenAI-compatible audio endpoints

vLLM Completions API

OpenAI-compatible text completions

vLLM Embeddings API

OpenAI-compatible embeddings

vLLM Scoring API

vLLM-specific scoring and reranking endpoints

vLLM Tokenize API

vLLM-specific tokenize/detokenize utilities

vLLM Chat Completions API

OpenAI-compatible chat completions

Scroll for all 7

Open Collections 8

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

API Collection

OPEN COLLECTION

Scroll for all 8

Pricing Plans 1

Published pricing tiers and plan structures.

Vllm Plans Pricing

1 plans

PLANS

Rate Limits 1

Documented rate limits and quota policies.

Vllm Rate Limits

2 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals for API financial operations.

Vllm Finops

FINOPS

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Vllm Authentication

http · 1 scheme

SECURITY

Vllm Domain Security

TLSv1.3

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Vllm Agentic Access

12 operations · 12 acting

12 operations · 12 acting

AGENTIC

Resources

Get Started 1

Portal, sign-up, and the first successful call

Agent Surfaces 2

MCP servers, agent skills, and machine-readable catalogs

Access & Security 2

Authentication, authorization, and security posture

Operate 1

Status, limits, changes, and where to get help

Commercial 2

Pricing, plans, and the legal terms of use

Company 3

The organization behind the API

Other 1

Properties that don't map to a standard resource type

Source (apis.yml)

apis.yml Raw ↑
aid: vllm
url: https://raw.githubusercontent.com/api-evangelist/vllm/refs/heads/main/apis.yml
name: vLLM
kind: company
description: vLLM is a high-throughput, memory-efficient open-source inference and serving engine for LLMs. It provides an
  OpenAI-compatible REST server (vllm serve) plus a Python API. vLLM is Apache 2.0 and run on your own GPU infrastructure;
  there is no hosted vLLM SaaS from the project itself.
deliveryModel:
  model: open-core
  license: Apache-2.0
  license_evidence:
    repo: vllm-project/vllm
    spdx: Apache-2.0
    stars: 92666
  open_source: true
  commercial: true
  callable_host: false
  label: Open core · an OSS project plus a commercial hosted product
  confidence: medium
  source:
  - openapi
  - pricing
  - repository-license
  generated: '2026-09-25'
  method: derived
accessModel:
  pricing: unknown
  onboarding: unknown
  trial: false
  try_now: false
  public: false
  label: Unknown
  confidence: low
  try_now_blocked_by: loopback
  source:
  - plans
  - authentication
  - security
  generated: '2026-09-03'
  method: derived
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/vllm.png
tags:
- LLM
- Inference
- Open Source
- GPU
- OpenAI-Compatible
- Self-Hosted
tags_raw:
- LLM
- Inference
- Open Source
- GPU
- OpenAI Compatible
- Self-Hosted
- Open-Source
created: '2026-05-08'
modified: '2026-05-08'
specificationVersion: '0.23'
apis:
- aid: vllm:openai-compatible
  name: vLLM OpenAI-Compatible Server
  description: OpenAI-compatible REST API exposed by `vllm serve`. Endpoints include /v1/chat/completions, /v1/completions,
    /v1/embeddings, /v1/score, /v1/audio/transcriptions, /v1/audio/translations, /v1/realtime (WebSocket), /tokenize, /detokenize,
    and /generative_scoring. Authentication via the --api-key flag set on server start; clients can use the official OpenAI
    Python library unmodified, with vLLM-specific extensions passed via extra_body.
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Chat
  - Completions
  - Embeddings
  - Audio
  - Score
  - OpenAI-Compatible
  properties:
  - type: Documentation
    url: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  - type: GitHub
    url: https://github.com/vllm-project/vllm
  - type: OpenAICompat
    url: https://platform.openai.com/docs/api-reference
- aid: vllm:vllm-audio-api
  name: vLLM Audio API
  description: OpenAI-compatible audio endpoints
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Audio
  properties:
  - type: OpenAPI
    url: openapi/vllm-audio-api-openapi.yml
- aid: vllm:vllm-completions-api
  name: vLLM Completions API
  description: OpenAI-compatible text completions
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Completions
  properties:
  - type: OpenAPI
    url: openapi/vllm-completions-api-openapi.yml
- aid: vllm:vllm-embeddings-api
  name: vLLM Embeddings API
  description: OpenAI-compatible embeddings
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Embeddings
  properties:
  - type: OpenAPI
    url: openapi/vllm-embeddings-api-openapi.yml
- aid: vllm:vllm-scoring-api
  name: vLLM Scoring API
  description: vLLM-specific scoring and reranking endpoints
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Scoring
  properties:
  - type: OpenAPI
    url: openapi/vllm-scoring-api-openapi.yml
- aid: vllm:vllm-tokenize-api
  name: vLLM Tokenize API
  description: vLLM-specific tokenize/detokenize utilities
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Tokenize
  properties:
  - type: OpenAPI
    url: openapi/vllm-tokenize-api-openapi.yml
- aid: vllm:vllm-chat-completions-api
  name: vLLM Chat Completions API
  description: OpenAI-compatible chat completions
  humanURL: https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html
  baseURL: http://localhost:8000/v1
  tags:
  - Chat Completions
  properties:
  - type: OpenAPI
    url: openapi/vllm-chat-completions-api-openapi.yml
common:
- type: AgenticAccess
  url: agentic-access/vllm-agentic-access.yml
- type: DomainSecurity
  url: security/vllm-domain-security.yml
- type: Authentication
  url: authentication/vllm-authentication.yml
- type: LinkedIn
  url: https://www.linkedin.com/company/vllm-project
- type: Website
  url: https://docs.vllm.ai/
- type: DeveloperPortal
  url: https://docs.vllm.ai/
- type: OpenSource
  url: https://github.com/vllm-project/vllm
- type: Plans
  url: plans/vllm-plans-pricing.yml
- type: RateLimits
  url: rate-limits/vllm-rate-limits.yml
- type: FinOps
  url: finops/vllm-finops.yml
- type: LlmsText
  url: https://vllm.ai/llms.txt
- url: https://vllm.ai/blog/rss.xml
  type: Blog
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com

Work with this as data

Every provider here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for providers

9 MCP tools reach this
  • find_providersBrowse and filter every provider in the catalog.
  • get_provider_artifactsEvery artifact this provider publishes, grouped by type.
  • get_provider_operationsEvery operation across all of their OpenAPIs — one call instead of parsing every spec.
  • get_provider_toolsEvery MCP tool they ship, with the operation each wraps.
  • get_provider_evidenceHow each part of their score was established. Free — the basis for a claim should not sit behind it.
  • get_provider_ratingPRO — composite, band, trend and facet scores.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This provider
curl "https://apis.io/api/v1/providers/vllm"
All providers
curl "https://apis.io/api/v1/providers?limit=25"
Every operation they expose
curl "https://apis.io/api/v1/providers/vllm/operations?limit=25"
How their score was established
curl "https://apis.io/api/v1/providers/vllm/evidence"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.