Vespa website screenshot

Vespa

Vespa is an open-source AI search engine, big-data serving engine, and vector database originally developed inside Yahoo and spun out as Vespa.ai AS. Vespa combines vector search, text search (BM25), structured filtering, and machine-learned ranking — including native tensor inference — into a single distributed serving engine that scales to billions of documents with sub-100ms latency. Vespa Cloud is the fully managed commercial offering operated by the Vespa.ai team across AWS and GCP, with Startup, Basic, Commercial, and Enterprise plans plus a Self-Managed option for customers running the open-source engine on their own infrastructure. Vespa is widely used at Spotify, Perplexity, Yahoo, Farfetch, and Elicit for search, recommendation, personalization, and Retrieval-Augmented Generation (RAG).

Vespa publishes 2 APIs on the APIs.io network: Document API and Query API. Tagged areas include AI, Search, Vector Database, Big Data, and Machine Learning.

The Vespa catalog on APIs.io includes 1 JSON-LD context and 2 Spectral governance rulesets.

Vespa’s developer surface includes authentication, documentation, getting-started guide, engineering blog, pricing, developer console, support, and 27 more developer resources.

60.4/100 strong ▬ flat Agent 31/100 agent aware Full breakdown ↓
scored 2026-08-05 · rubric v0.9.1
AccessFreemiumSelf serve⚡ Free to try
8 APIs 20 Features 6 Use Cases
AISearchVector DatabaseBig DataMachine LearningSemantic SearchRetrieval Augmented GenerationOpen SourceTensorRecommendations

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-08-05 · rubric v0.9.1
Composite quality — 60.4/100 · strong
Contract Quality 16.3 / 25
Developer Ergonomics 12.6 / 20
Commercial Clarity 10.0 / 20
Operational Transparency 6.8 / 13
Governance 8.3 / 12
Discoverability 6.5 / 10
Agent readiness — 31/100 · agent aware
Machine-Readable Contract 18 / 18
Agentic Access Contract 10 / 10
MCP Server 0 / 12
Machine-Readable Auth 10 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 0 / 7
Rate-Limit Signaling 7 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3
A2A Agent Card 0 / 8
Dry-Run / Simulate Mode 0 / 4
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. This rating is computed from github.com/api-evangelist/vespa-ai: open an issue to ask a question, or submit a pull request to add artifacts. Want it done for you? Prioritized profiling — $2,500 →

APIs 8

Individual APIs this provider publishes, each with its own machine-readable definition.

Vespa Document API

The Vespa Document API (/document/v1) provides synchronous REST access to document operations against a Vespa content cluster. It supports Put, Get, Update (partial update with ...

Vespa Deploy API

The Vespa Deploy API (/application/v2) manages application packages on a Vespa configuration server. It supports preparing, activating, and tearing down application packages, se...

Vespa Tenant and Application API

The Vespa Tenant API (/application/v2/tenant) manages tenants and applications hosted on a Vespa configuration server or Vespa Cloud control plane. It exposes operations for cre...

Vespa Config API

The Vespa Config API (/config/v2) lets services in a Vespa application retrieve their configuration from a Vespa configuration server using the config-server / config-proxy prot...

Vespa Cluster Controller API

The Vespa Cluster Controller API (/cluster/v2) exposes runtime state and management endpoints for a Vespa content cluster — including node state queries, maintenance-mode transi...

Vespa State API

The Vespa State API (/state/v1) exposes per-service health, version, and metrics endpoints for any Vespa node — used by orchestration tooling, monitoring agents, and load balanc...

Vespa Metrics API

Vespa exposes a family of metrics endpoints (/metrics/v1, /metrics/v2, /prometheus/v1) that publish Vespa engine and application metrics in JSON or Prometheus exposition format ...

Vespa Query API

The Query API from Vespa — 1 operation(s) for query.

Scroll for all 8

Postman Collections 1

Ready-to-run Postman collections for exercising this provider's APIs.

Open Collections 1

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

Vespa Query API

OPEN COLLECTION

Pricing Plans 1

Published pricing tiers and plan structures.

Rate Limits 1

Documented rate limits and quota policies.

Vespa Ai Rate Limits

6 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals for API financial operations.

Features 20

Notable capabilities this provider offers.

Open-source under Apache 2.0
Vector search with HNSW indexes
BM25 text search and hybrid search
Native tensor and ML model inference at serving time
YQL (Vespa Query Language) for structured queries
Multi-phase ranking (match-phase, first-phase, second-phase, global-phase)
Document API with conditional writes, visits, and JSON Lines streaming
Multi-tenant namespaces and document groups
Real-time indexing with sub-100ms query latency
Distributed content clusters with automatic sharding and replication
Streaming search mode for personal/private corpora
Built-in machine learning inference (TensorFlow, ONNX, XGBoost, LightGBM)
Approximate nearest neighbor and exact nearest neighbor operators
Application packages with schemas, services.xml, and rank profiles
Container API for custom searchers, document processors, and handlers
Self-managed (Apache 2.0) or fully managed Vespa Cloud (AWS, GCP)
Vespa Cloud Startup plan from $0.05 / vCPU-hour, $0.005 / GiB-memory-hour
Vespa Cloud Commercial plan with 24/7 1-hour SLA support
Vespa Cloud Enterprise plan with $20k/month minimum and 15-minute SLA
Up to 50% volume discounts and 15% committed-spend discount

Scroll for all 20

Semantic Vocabularies 1

JSON-LD contexts and semantic vocabularies used across these APIs.

Vespa Ai Context

11 classes · 6 properties

JSON-LD

Spectral Rules 2

Spectral governance rulesets for linting and validating these APIs.

Vespa API Rules

5 rules · 4 warnings 1 info

SPECTRAL

Vespa API Rules

7 rules · 7 warnings

SPECTRAL

JSON Schema 2

Standalone JSON Schema definitions for this provider's data models.

VespaDocument

3 properties

JSON SCHEMA

VespaQuery

10 properties

JSON SCHEMA

JSON Structure 1

JSON Structure definitions describing this provider's data shapes.

Vespa Ai Document Structure

0 properties

JSON STRUCTURE

Examples 2

Example request and response payloads for these APIs.

Vespa Ai Query Example

2 fields

EXAMPLE

Security Posture 3

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Vespa Ai Authentication

http · 1 scheme

SECURITY

Vespa Ai Domain Security

TLSv1.3 · HSTS · DMARC

SECURITY

Vespa Ai Vulnerability Disclosure

Intigriti · security.txt · contact published

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Vespa Ai Agentic Access

2 operations · 1 acting

2 operations · 1 acting

AGENTIC

Use Cases 6

What developers build with this provider.

Hybrid Search

Combine BM25 text relevance with vector similarity and structured filters in a single query executed by Vespa's multi-phase ranking pipeline.

Retrieval Augmented Generation

Serve grounded context to large language models by indexing documents, chunks, and embeddings in Vespa and retrieving them with hybrid search at sub-100ms latency.

Recommendation and Personalization

Power recommendation systems with machine-learned ranking, real-time feature updates, and tensor inference over user and item embeddings.

Ad Targeting and Real-Time Bidding

Match candidate ads against user context and serve ranked impressions within tight latency budgets using Vespa's distributed serving engine.

E-Commerce Search and Browse

Combine faceted navigation, structured filters, text relevance, and learned ranking for large product catalogs with frequent updates.

Streaming Search for Personal Data

Run "streaming search" mode that scans a user's personal corpus on demand — ideal for mail, messaging, and document search where each user has their own private index.

Integrations 12

Pre-built integrations with other platforms and tools.

AWS

Google Cloud

Prometheus

Grafana

TensorFlow

ONNX Runtime

XGBoost

LightGBM

Kubernetes

LangChain

LlamaIndex

Haystack

Scroll for all 12

Resources

Get Started 2

Portal, sign-up, and the first successful call

Documentation 1

Reference material describing how the API behaves

Agent Surfaces 1

MCP servers, agent skills, and machine-readable catalogs

Design & Contract 3

Pagination, idempotency, versioning, errors, and events

Build 10

SDKs, sample code, and the tooling you integrate with

Scroll for all 10

Access & Security 3

Authentication, authorization, and security posture

Learn 1

Tutorials, courses, talks, and written guidance

Operate 4

Status, limits, changes, and where to get help

Commercial 4

Pricing, plans, and the legal terms of use

Company 3

The organization behind the API

Other 2

Properties that don't map to a standard resource type

Source (apis.yml)

apis.yml Raw ↑
aid: vespa-ai
name: Vespa
description: Vespa is an open-source AI search engine, big-data serving engine, and vector database originally developed inside
  Yahoo and spun out as Vespa.ai AS. Vespa combines vector search, text search (BM25), structured filtering, and machine-learned
  ranking — including native tensor inference — into a single distributed serving engine that scales to billions of documents
  with sub-100ms latency. Vespa Cloud is the fully managed commercial offering operated by the Vespa.ai team across AWS and
  GCP, with Startup, Basic, Commercial, and Enterprise plans plus a Self-Managed option for customers running the open-source
  engine on their own infrastructure. Vespa is widely used at Spotify, Perplexity, Yahoo, Farfetch, and Elicit for search,
  recommendation, personalization, and Retrieval-Augmented Generation (RAG).
type: Index
accessModel:
  pricing: freemium
  onboarding: self-serve
  trial: false
  try_now: true
  public: false
  label: Freemium · Self-serve signup
  confidence: high
  source:
  - plans
  - authentication
  generated: '2026-07-22'
  method: derived
position: Provider
access: 3rd-Party
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/vespa-ai.png
tags:
- AI
- Search
- Vector Database
- Big Data
- Machine Learning
- Semantic Search
- Retrieval Augmented Generation
- Open Source
- Tensor
- Recommendations
url: https://raw.githubusercontent.com/api-evangelist/vespa-ai/refs/heads/main/apis.yml
created: '2026-05-25'
modified: '2026-05-25'
specificationVersion: '0.19'
apis:
- aid: vespa-ai:vespa-document-api
  name: Vespa Document API
  description: The Vespa Document API (/document/v1) provides synchronous REST access to document operations against a Vespa
    content cluster. It supports Put, Get, Update (partial update with assign/add/remove operators), Remove, and Visit (streaming
    visit, copy, delete-where, update-where) over JSON or JSON Lines, with conditional writes, multi-tenant namespaces, field-set
    projection, time-window selection, and pagination via continuation tokens.
  humanURL: https://docs.vespa.ai/en/reference/document-v1-api-reference.html
  tags:
  - Documents
  - CRUD
  - Indexing
  - Data
  - Streaming
  properties:
  - url: openapi/vespa-document-api-openapi.yml
    type: OpenAPI
  - url: https://docs.vespa.ai/en/reference/document-v1-api-reference.html
    type: Documentation
  - url: https://docs.vespa.ai/en/writing/document-v1-api-guide.html
    type: Documentation
  - url: https://docs.vespa.ai/en/reads-and-writes.html
    type: Documentation
- aid: vespa-ai:vespa-deploy-api
  name: Vespa Deploy API
  description: The Vespa Deploy API (/application/v2) manages application packages on a Vespa configuration server. It supports
    preparing, activating, and tearing down application packages, session-based deployments, schema validation, and zero-downtime
    updates of services, schemas, and rank profiles.
  humanURL: https://docs.vespa.ai/en/reference/deploy-rest-api-v2.html
  tags:
  - Deployment
  - Configuration
  - Application
  - DevOps
  properties:
  - url: https://docs.vespa.ai/en/reference/deploy-rest-api-v2.html
    type: Documentation
  - url: https://docs.vespa.ai/en/application-packages.html
    type: Documentation
- aid: vespa-ai:vespa-tenant-api
  name: Vespa Tenant and Application API
  description: The Vespa Tenant API (/application/v2/tenant) manages tenants and applications hosted on a Vespa configuration
    server or Vespa Cloud control plane. It exposes operations for creating tenants, listing applications, and binding application
    sessions to a tenant.
  humanURL: https://docs.vespa.ai/en/reference/application-v2-tenant.html
  tags:
  - Tenants
  - Applications
  - Multi-Tenancy
  - Administration
  properties:
  - url: https://docs.vespa.ai/en/reference/application-v2-tenant.html
    type: Documentation
- aid: vespa-ai:vespa-config-api
  name: Vespa Config API
  description: The Vespa Config API (/config/v2) lets services in a Vespa application retrieve their configuration from a
    Vespa configuration server using the config-server / config-proxy protocol. It is primarily used by Vespa services and
    tooling rather than end users, but is documented as a stable HTTP API.
  humanURL: https://docs.vespa.ai/en/reference/config-rest-api-v2.html
  tags:
  - Configuration
  - Internal
  properties:
  - url: https://docs.vespa.ai/en/reference/config-rest-api-v2.html
    type: Documentation
- aid: vespa-ai:vespa-cluster-api
  name: Vespa Cluster Controller API
  description: The Vespa Cluster Controller API (/cluster/v2) exposes runtime state and management endpoints for a Vespa content
    cluster — including node state queries, maintenance-mode transitions, and storage cluster orchestration.
  humanURL: https://docs.vespa.ai/en/reference/cluster-v2.html
  tags:
  - Cluster
  - Operations
  - Content
  - State
  properties:
  - url: https://docs.vespa.ai/en/reference/cluster-v2.html
    type: Documentation
- aid: vespa-ai:vespa-state-api
  name: Vespa State API
  description: The Vespa State API (/state/v1) exposes per-service health, version, and metrics endpoints for any Vespa node
    — used by orchestration tooling, monitoring agents, and load balancers to check liveness, readiness, and runtime metrics.
  humanURL: https://docs.vespa.ai/en/reference/state-v1.html
  tags:
  - Health
  - Monitoring
  - Metrics
  - Observability
  properties:
  - url: https://docs.vespa.ai/en/reference/state-v1.html
    type: Documentation
- aid: vespa-ai:vespa-metrics-api
  name: Vespa Metrics API
  description: Vespa exposes a family of metrics endpoints (/metrics/v1, /metrics/v2, /prometheus/v1) that publish Vespa engine
    and application metrics in JSON or Prometheus exposition format for scraping by Prometheus, Grafana, or other observability
    stacks.
  humanURL: https://docs.vespa.ai/en/operations/metrics.html
  tags:
  - Metrics
  - Prometheus
  - Observability
  - Monitoring
  properties:
  - url: https://docs.vespa.ai/en/operations/metrics.html
    type: Documentation
  - url: https://docs.vespa.ai/en/reference/metrics-v1.html
    type: Documentation
  - url: https://docs.vespa.ai/en/reference/metrics-v2.html
    type: Documentation
  - url: https://docs.vespa.ai/en/reference/prometheus-v1.html
    type: Documentation
- aid: vespa-ai:vespa-ai-query-api
  name: Vespa Query API
  description: The Query API from Vespa — 1 operation(s) for query.
  humanURL: https://docs.vespa.ai/en/query-api.html
  tags:
  - Query
  properties:
  - type: OpenAPI
    url: openapi/vespa-ai-query-api-openapi.yml
  - type: Documentation
    url: https://docs.vespa.ai/en/query-api.html
  - type: Documentation
    url: https://docs.vespa.ai/en/reference/api/query.html
  - type: GettingStarted
    url: https://docs.vespa.ai/en/getting-started.html
common:
- type: PostmanWorkspace
  url: https://www.postman.com/kinlaneapi/vespa/overview
- type: AgenticAccess
  url: agentic-access/vespa-ai-agentic-access.yml
- type: VulnerabilityDisclosure
  url: security/vespa-ai-vulnerability-disclosure.yml
- type: DomainSecurity
  url: security/vespa-ai-domain-security.yml
- type: Authentication
  url: authentication/vespa-ai-authentication.yml
- type: Website
  url: https://vespa.ai
- type: Documentation
  url: https://docs.vespa.ai/
- type: GettingStarted
  url: https://docs.vespa.ai/en/getting-started.html
- type: Tutorials
  url: https://docs.vespa.ai/en/learn/tutorials/
- type: GitHubOrganization
  url: https://github.com/vespa-engine
- type: GitHubRepository
  url: https://github.com/vespa-engine/vespa
- type: License
  url: https://github.com/vespa-engine/vespa/blob/master/LICENSE
- type: Blog
  url: https://blog.vespa.ai/
- type: BlogRSS
  url: https://blog.vespa.ai/feed.xml
- type: Pricing
  url: https://cloud.vespa.ai/pricing
- type: Console
  url: https://console.vespa-cloud.com/
- type: Slack
  url: https://slack.vespa.ai/
- type: Support
  url: https://github.com/vespa-engine/vespa/issues
- type: ChangeLog
  url: https://github.com/vespa-engine/vespa/releases
- type: SDKs
  name: Vespa CLI (Go)
  url: https://github.com/vespa-engine/vespa/tree/master/client/go
- type: SDKs
  name: pyvespa (Python)
  url: https://github.com/vespa-engine/pyvespa
- type: SDKs
  name: pyvespa Documentation
  url: https://vespa-engine.github.io/pyvespa/
- type: SDKs
  name: vespa-feed-client (Java)
  url: https://github.com/vespa-engine/vespa/tree/master/vespa-feed-client
- type: SDKs
  name: vespa-search (JavaScript)
  url: https://github.com/vespa-engine/vespa-search
- type: SampleApps
  url: https://github.com/vespa-engine/sample-apps
- type: PrometheusExporter
  url: https://github.com/vespa-engine/vespa_exporter
- type: DockerImage
  url: https://github.com/vespa-engine/docker-image
- type: GitHubAction
  url: https://github.com/vespa-engine/setup-vespa-cli-action
- type: SpectralRules
  url: rules/vespa-ai-rules.yml
- type: Vocabulary
  url: vocabulary/vespa-ai-vocabulary.yml
- type: JSONLDContext
  url: json-ld/vespa-ai-context.jsonld
- type: Plans
  url: plans/vespa-ai-plans-pricing.yml
- type: RateLimits
  url: rate-limits/vespa-ai-rate-limits.yml
- type: FinOps
  url: finops/vespa-ai-finops.yml
- type: Features
  data:
  - Open-source under Apache 2.0
  - Vector search with HNSW indexes
  - BM25 text search and hybrid search
  - Native tensor and ML model inference at serving time
  - YQL (Vespa Query Language) for structured queries
  - Multi-phase ranking (match-phase, first-phase, second-phase, global-phase)
  - Document API with conditional writes, visits, and JSON Lines streaming
  - Multi-tenant namespaces and document groups
  - Real-time indexing with sub-100ms query latency
  - Distributed content clusters with automatic sharding and replication
  - Streaming search mode for personal/private corpora
  - Built-in machine learning inference (TensorFlow, ONNX, XGBoost, LightGBM)
  - Approximate nearest neighbor and exact nearest neighbor operators
  - Application packages with schemas, services.xml, and rank profiles
  - Container API for custom searchers, document processors, and handlers
  - Self-managed (Apache 2.0) or fully managed Vespa Cloud (AWS, GCP)
  - Vespa Cloud Startup plan from $0.05 / vCPU-hour, $0.005 / GiB-memory-hour
  - Vespa Cloud Commercial plan with 24/7 1-hour SLA support
  - Vespa Cloud Enterprise plan with $20k/month minimum and 15-minute SLA
  - Up to 50% volume discounts and 15% committed-spend discount
  sources:
  - https://cloud.vespa.ai/price-calculator.html
  - https://docs.vespa.ai/
  updated: '2026-05-25'
- type: UseCases
  data:
  - name: Hybrid Search
    description: Combine BM25 text relevance with vector similarity and structured filters in a single query executed by Vespa's
      multi-phase ranking pipeline.
  - name: Retrieval Augmented Generation
    description: Serve grounded context to large language models by indexing documents, chunks, and embeddings in Vespa and
      retrieving them with hybrid search at sub-100ms latency.
  - name: Recommendation and Personalization
    description: Power recommendation systems with machine-learned ranking, real-time feature updates, and tensor inference
      over user and item embeddings.
  - name: Ad Targeting and Real-Time Bidding
    description: Match candidate ads against user context and serve ranked impressions within tight latency budgets using
      Vespa's distributed serving engine.
  - name: E-Commerce Search and Browse
    description: Combine faceted navigation, structured filters, text relevance, and learned ranking for large product catalogs
      with frequent updates.
  - name: Streaming Search for Personal Data
    description: Run "streaming search" mode that scans a user's personal corpus on demand — ideal for mail, messaging, and
      document search where each user has their own private index.
- type: Integrations
  data:
  - name: AWS
  - name: Google Cloud
  - name: Prometheus
  - name: Grafana
  - name: TensorFlow
  - name: ONNX Runtime
  - name: XGBoost
  - name: LightGBM
  - name: Kubernetes
  - name: LangChain
  - name: LlamaIndex
  - name: Haystack
integrations:
- name: AWS
- name: Google Cloud
- name: LangChain
- name: LlamaIndex
- name: Haystack
- name: Prometheus
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com