# vLLM OpenAI-Compatible API

**Canonical:** https://apis.io/apis/scalable-inference-serving/vllm-openai-compatible-api/  
**Provider:** Scalable Inference Serving — https://apis.io/providers/scalable-inference-serving/

vLLM OpenAI-Compatible API is one of 9 APIs that [Scalable Inference Serving](https://apis.io/providers/scalable-inference-serving/) publishes on the [APIs.io](https://apis.io/) network. Tagged areas include GPU, Inference, KV Cache, LLM, and Model Serving. The published artifact set on APIs.io includes API documentation, an API reference, and a changelog.

vLLM is a high-throughput and memory-efficient inference engine for LLMs, implementing PagedAttention for efficient KV cache management. vLLM exposes an OpenAI-compatible REST API allowing seamless migration from OpenAI endpoints. In 2026, vLLM integrates with KServe via LLMInferenceService and llm-d for production-grade distributed LLM inference. Powers major LLM deployments at scale.

## Machine-readable artifacts (4)

- **Documentation** — https://docs.vllm.ai/en/stable/
- **GitHub** — https://github.com/vllm-project/vllm
- **APIReference** — https://docs.vllm.ai/en/stable/serving/openai_compatible_server.html
- **ChangeLog** — https://github.com/vllm-project/vllm/releases

## Other Scalable Inference Serving APIs (8)

- [BentoML REST API](https://apis.io/apis/scalable-inference-serving/bentoml-rest-api/)
- [NVIDIA Triton Inference Server HTTP API](https://apis.io/apis/scalable-inference-serving/nvidia-triton-inference-server-http-api/)
- [MLflow Model Registry REST API](https://apis.io/apis/scalable-inference-serving/mlflow-model-registry-rest-api/)
- [Ray Serve REST API](https://apis.io/apis/scalable-inference-serving/ray-serve-rest-api/)
- [Scalable Inference Serving Health API](https://apis.io/apis/scalable-inference-serving/scalable-inference-serving-health-api/)
- [Scalable Inference Serving Inference API](https://apis.io/apis/scalable-inference-serving/scalable-inference-serving-inference-api/)
- [Scalable Inference Serving Metadata API](https://apis.io/apis/scalable-inference-serving/scalable-inference-serving-metadata-api/)
- [Scalable Inference Serving Models API](https://apis.io/apis/scalable-inference-serving/scalable-inference-serving-models-api/)

## Tags

GPU, Inference, KV Cache, LLM, Model Serving, Open-Source, OpenAI-Compatible

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/apis/scalable-inference-serving/vllm-openai-compatible-api/). The API's provider profile, Kin Score and agent-readiness rating are at https://apis.io/providers/scalable-inference-serving/.
