# BentoML REST API

**Canonical:** https://apis.io/apis/scalable-inference-serving/bentoml-rest-api/  
**Provider:** Scalable Inference Serving — https://apis.io/providers/scalable-inference-serving/

BentoML REST API is one of 9 APIs that [Scalable Inference Serving](https://apis.io/providers/scalable-inference-serving/) publishes on the [APIs.io](https://apis.io/) network. Tagged areas include Batching, Inference, Model Serving, Open-Source, and Python. The published artifact set on APIs.io includes API documentation, a getting-started guide, pricing, and an API reference.

BentoML is an open-source unified inference platform for deploying and scaling AI models. It auto-generates RESTful APIs from Python service definitions, provides built-in OpenAPI/Swagger documentation, supports adaptive batching, and integrates with KServe for Kubernetes deployment. BentoML 1.0 introduced the Runner abstraction for parallelizing inference workloads with adaptive batching and independent scaling of pre/post-processing from model inference.

## Machine-readable artifacts (6)

- **Documentation** — https://docs.bentoml.com/en/latest/
- **GitHub** — https://github.com/bentoml/BentoML
- **GettingStarted** — https://docs.bentoml.com/en/latest/get-started/quickstart.html
- **Pricing** — https://www.bentoml.com/pricing
- **APIReference** — https://docs.bentoml.com/en/latest/reference/index.html
- **APIsJSON** — https://raw.githubusercontent.com/api-evangelist/scalable-inference-serving/refs/heads/main/apis.yml

## Other Scalable Inference Serving APIs (8)

- [vLLM OpenAI-Compatible API](https://apis.io/apis/scalable-inference-serving/vllm-openai-compatible-api/)
- [NVIDIA Triton Inference Server HTTP API](https://apis.io/apis/scalable-inference-serving/nvidia-triton-inference-server-http-api/)
- [MLflow Model Registry REST API](https://apis.io/apis/scalable-inference-serving/mlflow-model-registry-rest-api/)
- [Ray Serve REST API](https://apis.io/apis/scalable-inference-serving/ray-serve-rest-api/)
- [Scalable Inference Serving Health API](https://apis.io/apis/scalable-inference-serving/scalable-inference-serving-health-api/)
- [Scalable Inference Serving Inference API](https://apis.io/apis/scalable-inference-serving/scalable-inference-serving-inference-api/)
- [Scalable Inference Serving Metadata API](https://apis.io/apis/scalable-inference-serving/scalable-inference-serving-metadata-api/)
- [Scalable Inference Serving Models API](https://apis.io/apis/scalable-inference-serving/scalable-inference-serving-models-api/)

## Tags

Batching, Inference, Model Serving, Open-Source, Python, REST API

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/apis/scalable-inference-serving/bentoml-rest-api/). The API's provider profile, Kin Score and agent-readiness rating are at https://apis.io/providers/scalable-inference-serving/.
