AI Gateway website screenshot

AI Gateway

An API Evangelist landscape index of AI gateways — the LLM routers, prompt firewalls, model fallback proxies, cost-control planes, and policy engines that sit between applications and AI providers. AI gateways unify access across OpenAI, Anthropic, Google, AWS Bedrock, Azure OpenAI, and self-hosted models behind a common interface and apply caching, routing, guardrails, observability, rate limiting, budgets, RBAC, and audit controls. This index catalogs commercial SaaS gateways, open-source projects, API gateway AI plugins, and cloud-provider AI proxies, with a shared schema and vocabulary for describing model routes, fallbacks, guardrails, and budgets across vendors.

AI Gateway publishes 23 APIs on the APIs.io network, including Analytics API, APIKeys API, Assistants API, and 20 more. Tagged areas include AI Gateway, LLM Router, LLM Proxy, Model Routing, and Prompt Firewall.

The AI Gateway catalog on APIs.io includes 1 JSON-LD context and 1 Spectral governance ruleset.

AI Gateway’s developer surface includes authentication, code examples, developer portal, engineering blog, and 10 more developer resources.

33.9/100 thin ▬ flat Agent 29/100 agent aware Full breakdown ↓
scored 2026-07-28 · rubric v0.6
AccessSelf serve
38 APIs 13 Features 6 Use Cases
AI GatewayLLM RouterLLM ProxyModel RoutingPrompt FirewallGuardrailsAI ObservabilityCost ControlsAI GovernanceAPI Gateway

Kin Score

Kin Score Kin Score How this is scored →
scored 2026-07-28 · rubric v0.6
Composite quality — 33.9/100 · thin
Contract Quality 14.8 / 25
Developer Ergonomics 4.3 / 20
Commercial Clarity 0.0 / 20
Operational Transparency 0.0 / 13
Governance 8.3 / 12
Discoverability 6.5 / 10
Agent readiness — 29/100 · agent aware
Machine-Readable Contract 18 / 18
Agentic Access Contract 10 / 10
MCP Server 0 / 12
Machine-Readable Auth 10 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 7 / 7
Rate-Limit Signaling 0 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3
A2A Agent Card 0 / 8
Dry-Run / Simulate Mode 0 / 4
Improve this rating by publishing the missing artifacts — every area above can be raised, and the full rubric is at apis.io/rating/. This rating is computed from github.com/api-evangelist/ai-gateway: open an issue to ask a question, or submit a pull request to add artifacts. Want it done for you? Prioritized profiling — $2,500 →

APIs 38

Individual APIs this provider publishes, each with its own machine-readable definition.

Portkey

Portkey is a production-grade AI gateway and control plane that fronts 1,600+ LLMs with unified routing, fallbacks, semantic caching, guardrails, cost attribution, and prompt ma...

OpenRouter

OpenRouter is a unified inference marketplace exposing 400+ models from 60+ providers behind one OpenAI-compatible API, with automatic provider fallback, pay-as-you-go credits, ...

LiteLLM

LiteLLM (BerriAI) is an open-source LLM gateway that exposes 100+ LLM providers — OpenAI, Anthropic, Azure, Bedrock, Gemini — through a single OpenAI-compatible API. The LiteLLM...

Helicone

Helicone is an open-source AI observability and routing platform centered on requests, sessions, prompts, datasets, rate limits, and alerts. Integrates with OpenAI, Anthropic, G...

Cloudflare AI Gateway

Cloudflare AI Gateway is an edge-deployed proxy that fronts AI providers — Workers AI, Anthropic, Google Gemini, OpenAI, Replicate, and more — with caching, rate limiting, analy...

Kong AI Gateway

The Kong AI Gateway is delivered as the AI Proxy plugin for Kong Gateway, transforming and proxying requests across 16+ providers including OpenAI, Azure OpenAI, Anthropic, Amaz...

Apache APISIX AI Proxy

The Apache APISIX ai-proxy plugin streamlines integration with LLMs by converting plugin settings into the appropriate request format for OpenAI, DeepSeek, Azure OpenAI, Anthrop...

Tetrate Agent Router Service

Tetrate Agent Router Service is an Envoy AI Gateway-as-a-service from the creators of Envoy, providing an approved LLM catalog, unified model access, automatic fallback, cost ma...

NVIDIA NIM

NVIDIA NIM is a set of inference microservices for streamlined AI model deployment, prebuilt and optimized for low-latency, high-throughput inference on NVIDIA-accelerated infra...

Traefik AI Gateway

Traefik AI Gateway is an enterprise, self-hosted, Kubernetes-native AI gateway with safety and governance (NVIDIA Safety NIMs, jailbreak detection, content filtering across 22+ ...

Together AI

Together AI is a full-stack AI Native Cloud for inference, fine-tuning, and GPU clusters powered by research, exposing serverless inference, batch processing, dedicated model an...

Anyscale

Anyscale is the production-scale AI platform built on Ray by the creators of Ray, supporting LLM inference and other data-intensive AI workloads across distributed GPU clusters....

LangDB

LangDB is an enterprise AI gateway for routing and governing LLM traffic across providers, with observability, cost tracking, and policy enforcement. Public homepage was unreach...

Envoy AI Gateway

Envoy AI Gateway is an open-source extension to Envoy Proxy and Envoy Gateway, providing a Kubernetes-native AI traffic plane for routing, governing, and observing LLM calls acr...

Gentrace

Gentrace was an AI evaluation and observability product; the company has shut down and its codebase is now MIT-licensed open source on GitHub. Included here for historical compl...

AI Gateway Analytics API

The Analytics API from AI Gateway — 2 operation(s) for analytics.

AI Gateway APIKeys API

The APIKeys API from AI Gateway — 1 operation(s) for apikeys.

AI Gateway Assistants API

The Assistants API from AI Gateway — 1 operation(s) for assistants.

AI Gateway Audio API

The Audio API from AI Gateway — 3 operation(s) for audio.

AI Gateway Batches API

The Batches API from AI Gateway — 2 operation(s) for batches.

AI Gateway Chat API

The Chat API from AI Gateway — 1 operation(s) for chat.

AI Gateway Completions API

The Completions API from AI Gateway — 1 operation(s) for completions.

AI Gateway Configs API

The Configs API from AI Gateway — 1 operation(s) for configs.

AI Gateway Embeddings API

The Embeddings API from AI Gateway — 2 operation(s) for embeddings.

AI Gateway Feedback API

The Feedback API from AI Gateway — 1 operation(s) for feedback.

AI Gateway Files API

The Files API from AI Gateway — 2 operation(s) for files.

AI Gateway FineTuning API

The FineTuning API from AI Gateway — 2 operation(s) for finetuning.

AI Gateway Guardrails API

The Guardrails API from AI Gateway — 1 operation(s) for guardrails.

AI Gateway Images API

The Images API from AI Gateway — 1 operation(s) for images.

AI Gateway Integrations API

The Integrations API from AI Gateway — 2 operation(s) for integrations.

AI Gateway Logs API

The Logs API from AI Gateway — 2 operation(s) for logs.

AI Gateway MCP API

The MCP API from AI Gateway — 2 operation(s) for mcp.

AI Gateway Policies API

The Policies API from AI Gateway — 2 operation(s) for policies.

AI Gateway Prompts API

The Prompts API from AI Gateway — 6 operation(s) for prompts.

AI Gateway Responses API

The Responses API from AI Gateway — 1 operation(s) for responses.

AI Gateway Threads API

The Threads API from AI Gateway — 3 operation(s) for threads.

AI Gateway VirtualKeys API

The VirtualKeys API from AI Gateway — 1 operation(s) for virtualkeys.

AI Gateway Workspaces API

The Workspaces API from AI Gateway — 4 operation(s) for workspaces.

Scroll for all 38

Open Collections 1

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

Portkey AI Gateway API

OPEN COLLECTION

Features 13

Notable capabilities this provider offers.

Provider Abstraction

A unified, typically OpenAI-compatible API surface that lets clients call any supported LLM provider without provider-specific SDK juggling.

Model Routing

Route requests to the right model and provider based on alias, header, request content, identity, time-of-day, cost, or latency.

Fallback and Failover

Automatically retry failed requests against backup providers or models when a primary upstream is degraded, rate-limited, or down.

Load Balancing and Fanout

Distribute traffic across multiple providers or replicas using weighted, priority-based, or RPM/TPM-aware load balancing.

Response Caching

Exact-match and semantic caching of model responses to cut latency and provider spend; some gateways claim 40-70 percent cost savings.

Cost Controls and Budgets

Per-user, per-team, per-key, per-project budgets, spend tracking, and hard or soft caps on token consumption.

Rate Limiting and Quotas

RPM, TPM, concurrency, and per-key quotas enforced at the gateway, decoupled from each upstream provider's limits.

Guardrails and Prompt Firewall

Prompt injection detection, jailbreak filtering, content moderation, PII redaction, and topic control applied to requests and responses.

Observability

Request, response, token, cost, latency, error, and trace data exported via OpenTelemetry, Langfuse, Phoenix, Langsmith, or built-in dashboards.

Authentication and RBAC

Virtual keys, JWT, OAuth2, SSO, and role-based access control over which clients can use which models with which budgets.

BYOK and Secret Management

Bring-your-own provider API keys, with the gateway holding and injecting them so clients never see upstream credentials.

Multi-Tenant Governance

Per-tenant isolation of keys, budgets, logs, and policies for platform teams serving multiple internal product teams.

MCP Federation

Some AI gateways also front Model Context Protocol servers, aggregating tools and exposing a single MCP endpoint to agents.

Scroll for all 13

Semantic Vocabularies 1

JSON-LD contexts and semantic vocabularies used across these APIs.

Ai Gateway Context

9 classes · 70 properties

JSON-LD

Spectral Rules 1

Spectral governance rulesets for linting and validating these APIs.

AI Gateway API Rules

5 rules · 3 warnings 2 info

SPECTRAL

JSON Schema 3

Standalone JSON Schema definitions for this provider's data models.

AIGatewayPolicy

12 properties

JSON SCHEMA

AIGatewayProvider

10 properties

JSON SCHEMA

AIGatewayRoute

12 properties

JSON SCHEMA

JSON Structure 3

JSON Structure definitions describing this provider's data shapes.

Ai Gateway Policy Structure

12 properties

JSON STRUCTURE

Ai Gateway Provider Structure

9 properties

JSON STRUCTURE

Ai Gateway Route Structure

10 properties

JSON STRUCTURE

Examples 6

Example request and response payloads for these APIs.

Ai Gateway Provider Example

10 fields

EXAMPLE

Ai Gateway Route Example

12 fields

EXAMPLE

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Ai Gateway Authentication

apiKey · 2 schemes

SECURITY

Ai Gateway Domain Security

TLSv1.3 · HSTS · DNSSEC · DMARC

SECURITY

Agentic Access 1

Recommended x-agentic-access execution contracts for AI agents.

Ai Gateway Agentic Access

60 operations · 32 acting

60 operations · 32 acting

AGENTIC

Use Cases 6

What developers build with this provider.

Provider-Agnostic LLM Access

Front many LLM providers behind one API so application teams can switch models without changing client code.

Cost Containment for AI

Apply caching, routing to cheaper models, and per-team budgets to keep generative-AI spend predictable.

Reliability and Failover

Survive single-provider outages by automatically failing over to backup models when the primary degrades.

Centralized AI Governance

Enforce content, PII, and policy controls in one place for every AI request leaving the organization.

Observability and FinOps

Attribute cost and latency to teams, projects, and users; expose token-level metrics to FinOps and SRE.

Multi-Tenant AI Platforms

Build internal AI platforms where each product team gets its own virtual keys, budgets, and logs.

Integrations 9

Pre-built integrations with other platforms and tools.

OpenAI

Front OpenAI's GPT, embeddings, and image models behind the gateway with virtual keys and budgets.

Anthropic

Route Claude requests through the gateway for fallback, caching, and central observability.

Google Gemini and Vertex AI

Proxy Google Gemini and Vertex AI calls with OpenAI-format translation where supported.

AWS Bedrock

Bridge OpenAI-format clients to Bedrock-hosted Anthropic, Mistral, Cohere, Meta, and Amazon models.

Azure OpenAI

Route to Azure-hosted OpenAI deployments with per-region failover and key rotation.

Ollama and vLLM

Front self-hosted Ollama and vLLM inference servers for hybrid cloud and on-prem inference.

OpenTelemetry

Export request, token, cost, and trace data to any OTel-compatible observability backend.

Langfuse and Phoenix

Stream prompts, completions, and evaluations to Langfuse and Arize Phoenix for prompt and model analytics.

Model Context Protocol

Some AI gateways federate MCP servers alongside LLM routes, exposing a unified agent endpoint.

Scroll for all 9

Resources

Get Started 1

Portal, sign-up, and the first successful call

Documentation 3

Reference material describing how the API behaves

Agent Surfaces 1

MCP servers, agent skills, and machine-readable catalogs

Design & Contract 5

Pagination, idempotency, versioning, errors, and events

Build 1

SDKs, sample code, and the tooling you integrate with

Access & Security 2

Authentication, authorization, and security posture

Company 1

The organization behind the API

Source (apis.yml)

apis.yml Raw ↑
aid: ai-gateway
name: AI Gateway
description: An API Evangelist landscape index of AI gateways — the LLM routers, prompt firewalls, model fallback proxies,
  cost-control planes, and policy engines that sit between applications and AI providers. AI gateways unify access across
  OpenAI, Anthropic, Google, AWS Bedrock, Azure OpenAI, and self-hosted models behind a common interface and apply caching,
  routing, guardrails, observability, rate limiting, budgets, RBAC, and audit controls. This index catalogs commercial SaaS
  gateways, open-source projects, API gateway AI plugins, and cloud-provider AI proxies, with a shared schema and vocabulary
  for describing model routes, fallbacks, guardrails, and budgets across vendors.
url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/apis.yml
humanURL: https://github.com/api-evangelist/ai-gateway
type: Index
accessModel:
  pricing: unknown
  onboarding: self-serve
  trial: false
  try_now: false
  public: false
  label: Self-serve signup
  confidence: medium
  source:
  - authentication
  generated: '2026-07-22'
  method: derived
position: Consuming
access: 3rd-Party
image: https://kinlane-images.s3.amazonaws.com/shared/apis-json/icons/ai-gateway.png
tags:
- AI Gateway
- LLM Router
- LLM Proxy
- Model Routing
- Prompt Firewall
- Guardrails
- AI Observability
- Cost Controls
- AI Governance
- API Gateway
created: '2026-05-22'
modified: '2026-05-22'
specificationVersion: '0.19'
apis:
- aid: ai-gateway:portkey
  name: Portkey
  description: Portkey is a production-grade AI gateway and control plane that fronts 1,600+ LLMs with unified routing, fallbacks,
    semantic caching, guardrails, cost attribution, and prompt management. The open-source Portkey Gateway is MIT-licensed;
    a hosted SaaS adds governance, observability, and enterprise controls.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - AI Gateway
  - LLM Router
  - Guardrails
  - Observability
  - Prompt Management
  - Open Source
  properties:
  - type: Portal
    url: https://portkey.ai/
  - type: Documentation
    url: https://portkey.ai/docs/
  - type: GitHubRepository
    url: https://github.com/Portkey-AI/gateway
  - type: GitHubOrganization
    url: https://github.com/Portkey-AI
  x-deployment:
  - cloud
  - self-host
  - opensource
  x-license: MIT
- aid: ai-gateway:openrouter
  name: OpenRouter
  description: OpenRouter is a unified inference marketplace exposing 400+ models from 60+ providers behind one OpenAI-compatible
    API, with automatic provider fallback, pay-as-you-go credits, custom data policies, and edge-routed latency optimization.
    It is a proprietary SaaS service.
  humanURL: https://openrouter.ai/
  baseURL: https://openrouter.ai/api/v1
  tags:
  - AI Gateway
  - LLM Marketplace
  - Multi-Provider
  - Fallback
  - Proprietary
  properties:
  - type: Portal
    url: https://openrouter.ai/
  - type: Documentation
    url: https://openrouter.ai/docs
  - type: Models
    url: https://openrouter.ai/models
  x-deployment:
  - cloud
  x-license: Proprietary
- aid: ai-gateway:litellm
  name: LiteLLM
  description: LiteLLM (BerriAI) is an open-source LLM gateway that exposes 100+ LLM providers — OpenAI, Anthropic, Azure,
    Bedrock, Gemini — through a single OpenAI-compatible API. The LiteLLM Proxy adds virtual keys, load balancing, RPM/TPM
    limits, spend tracking, and observability hooks for Langfuse, Phoenix, Langsmith, and OpenTelemetry. Self-hostable via
    Docker; enterprise support available.
  humanURL: https://www.litellm.ai/
  baseURL: https://api.litellm.ai
  tags:
  - AI Gateway
  - LLM Proxy
  - Open Source
  - Cost Tracking
  - Load Balancing
  properties:
  - type: Portal
    url: https://www.litellm.ai/
  - type: Documentation
    url: https://docs.litellm.ai/
  - type: GitHubRepository
    url: https://github.com/BerriAI/litellm
  - type: PyPI
    url: https://pypi.org/project/litellm/
  x-deployment:
  - self-host
  - opensource
  - cloud
  x-license: MIT
- aid: ai-gateway:helicone
  name: Helicone
  description: Helicone is an open-source AI observability and routing platform centered on requests, sessions, prompts, datasets,
    rate limits, and alerts. Integrates with OpenAI, Anthropic, Google Gemini, DeepSeek, Together AI, Mistral, Groq, Azure,
    OpenRouter, and LiteLLM. Available as managed cloud or self-hosted.
  humanURL: https://www.helicone.ai/
  baseURL: https://api.helicone.ai
  tags:
  - AI Gateway
  - Observability
  - Prompt Management
  - Open Source
  - Caching
  properties:
  - type: Portal
    url: https://www.helicone.ai/
  - type: Documentation
    url: https://docs.helicone.ai/
  - type: GitHubRepository
    url: https://github.com/Helicone/helicone
  x-deployment:
  - cloud
  - self-host
  - opensource
- aid: ai-gateway:cloudflare-ai-gateway
  name: Cloudflare AI Gateway
  description: Cloudflare AI Gateway is an edge-deployed proxy that fronts AI providers — Workers AI, Anthropic, Google Gemini,
    OpenAI, Replicate, and more — with caching, rate limiting, analytics, and request logging. Available on all Cloudflare
    plans.
  humanURL: https://developers.cloudflare.com/ai-gateway/
  baseURL: https://gateway.ai.cloudflare.com
  tags:
  - AI Gateway
  - Edge
  - Caching
  - Rate Limiting
  - Analytics
  properties:
  - type: Portal
    url: https://www.cloudflare.com/developer-platform/ai-gateway/
  - type: Documentation
    url: https://developers.cloudflare.com/ai-gateway/
  - type: GettingStarted
    url: https://developers.cloudflare.com/ai-gateway/get-started/
  x-deployment:
  - cloud
  x-license: Proprietary
- aid: ai-gateway:kong-ai-gateway
  name: Kong AI Gateway
  description: The Kong AI Gateway is delivered as the AI Proxy plugin for Kong Gateway, transforming and proxying requests
    across 16+ providers including OpenAI, Azure OpenAI, Anthropic, Amazon Bedrock, Gemini, Vertex AI, Cohere, Mistral, Hugging
    Face, Llama, xAI, Ollama, Alibaba DashScope, Cerebras, DeepSeek, Databricks, and vLLM. Supports chat, completions, embeddings,
    assistants, audio, image, video, batches, and files routes with template-based model selection.
  humanURL: https://konghq.com/products/kong-ai-gateway
  baseURL: https://konghq.com
  tags:
  - AI Gateway
  - API Gateway
  - Multi-Provider
  - Plugin
  - Kong
  properties:
  - type: Portal
    url: https://konghq.com/products/kong-ai-gateway
  - type: Documentation
    url: https://developer.konghq.com/plugins/ai-proxy/
  - type: GitHubOrganization
    url: https://github.com/Kong
  x-deployment:
  - cloud
  - self-host
  - opensource
- aid: ai-gateway:apisix-ai-proxy
  name: Apache APISIX AI Proxy
  description: The Apache APISIX ai-proxy plugin streamlines integration with LLMs by converting plugin settings into the
    appropriate request format for OpenAI, DeepSeek, Azure OpenAI, Anthropic, Google Gemini, Vertex AI, OpenRouter, AIMLAPI,
    and OpenAI-compatible services. Supports embedding models, observability of token usage and latency, custom endpoints,
    and flexible authentication. Apache 2.0 licensed.
  humanURL: https://apisix.apache.org/
  baseURL: https://apisix.apache.org
  tags:
  - AI Gateway
  - API Gateway
  - Open Source
  - Apache
  - Plugin
  properties:
  - type: Portal
    url: https://apisix.apache.org/
  - type: Documentation
    url: https://apisix.apache.org/docs/apisix/plugins/ai-proxy/
  - type: GitHubRepository
    url: https://github.com/apache/apisix
  x-deployment:
  - self-host
  - opensource
  x-license: Apache-2.0
- aid: ai-gateway:tetrate-agent-router
  name: Tetrate Agent Router Service
  description: Tetrate Agent Router Service is an Envoy AI Gateway-as-a-service from the creators of Envoy, providing an approved
    LLM catalog, unified model access, automatic fallback, cost management, AI guardrails, and an MCP gateway for agent tool
    connectivity. Built on Envoy AI Gateway.
  humanURL: https://tetrate.io/products/tetrate-agent-router-service/
  baseURL: https://tetrate.io
  tags:
  - AI Gateway
  - Envoy
  - MCP Gateway
  - Guardrails
  - Multi-Provider
  properties:
  - type: Portal
    url: https://tetrate.io/products/tetrate-agent-router-service/
  - type: Documentation
    url: https://docs.tetrate.io/
  - type: GitHubOrganization
    url: https://github.com/envoyproxy
  - type: GitHubRepository
    url: https://github.com/envoyproxy/ai-gateway
  x-deployment:
  - cloud
  - self-host
  - opensource
- aid: ai-gateway:nvidia-nim
  name: NVIDIA NIM
  description: NVIDIA NIM is a set of inference microservices for streamlined AI model deployment, prebuilt and optimized
    for low-latency, high-throughput inference on NVIDIA-accelerated infrastructure. Includes TensorRT and TensorRT-LLM-backed
    engines and exposes stable OpenAI-compatible APIs for self-hosted and cloud deployment.
  humanURL: https://www.nvidia.com/en-us/ai/
  baseURL: https://build.nvidia.com
  tags:
  - AI Gateway
  - Inference
  - Self-Hosted
  - NVIDIA
  - GPU
  properties:
  - type: Portal
    url: https://build.nvidia.com/
  - type: Documentation
    url: https://docs.nvidia.com/nim/
  - type: GitHubOrganization
    url: https://github.com/NVIDIA
  x-deployment:
  - self-host
  - cloud
  x-license: Proprietary
- aid: ai-gateway:traefik-ai-gateway
  name: Traefik AI Gateway
  description: Traefik AI Gateway is an enterprise, self-hosted, Kubernetes-native AI gateway with safety and governance (NVIDIA
    Safety NIMs, jailbreak detection, content filtering across 22+ categories), multi-LLM support via an OpenAI-compatible
    interface (Anthropic, Azure OpenAI, AWS Bedrock, Cohere, Gemini, Mistral, Ollama), intelligent routing, credential management,
    semantic caching with claimed 40-70 percent cost savings, PII protection via Presidio (35+ recognizers), and OpenTelemetry
    observability.
  humanURL: https://traefik.io/solutions/ai-gateway/
  baseURL: https://traefik.io
  tags:
  - AI Gateway
  - Kubernetes
  - Guardrails
  - Semantic Caching
  - PII Protection
  properties:
  - type: Portal
    url: https://traefik.io/solutions/ai-gateway/
  - type: Documentation
    url: https://doc.traefik.io/
  - type: GitHubOrganization
    url: https://github.com/traefik
  x-deployment:
  - self-host
  - cloud
- aid: ai-gateway:together-ai
  name: Together AI
  description: Together AI is a full-stack AI Native Cloud for inference, fine-tuning, and GPU clusters powered by research,
    exposing serverless inference, batch processing, dedicated model and container inference, GPU clusters, fine-tuning, managed
    storage, and code sandboxes for open-source models.
  humanURL: https://www.together.ai/
  baseURL: https://api.together.xyz
  tags:
  - Inference
  - Open Models
  - GPU
  - Multi-Provider
  - SaaS
  properties:
  - type: Portal
    url: https://www.together.ai/
  - type: Documentation
    url: https://docs.together.ai/
  - type: GitHubOrganization
    url: https://github.com/togethercomputer
  x-deployment:
  - cloud
  x-license: Proprietary
- aid: ai-gateway:anyscale
  name: Anyscale
  description: Anyscale is the production-scale AI platform built on Ray by the creators of Ray, supporting LLM inference
    and other data-intensive AI workloads across distributed GPU clusters. Integrates with vLLM and SkyRL; users bring their
    own models.
  humanURL: https://www.anyscale.com/
  baseURL: https://api.endpoints.anyscale.com
  tags:
  - Inference
  - Ray
  - GPU
  - Open Source
  - Self-Hosted
  properties:
  - type: Portal
    url: https://www.anyscale.com/
  - type: Documentation
    url: https://docs.anyscale.com/
  - type: GitHubOrganization
    url: https://github.com/anyscale
  - type: GitHubRepository
    url: https://github.com/ray-project/ray
  x-deployment:
  - cloud
  - self-host
- aid: ai-gateway:langdb
  name: LangDB
  description: LangDB is an enterprise AI gateway for routing and governing LLM traffic across providers, with observability,
    cost tracking, and policy enforcement. Public homepage was unreachable for direct verification during this profiling pass;
    see GitHub for current capabilities.
  humanURL: https://www.langdb.ai/
  baseURL: https://api.langdb.ai
  tags:
  - AI Gateway
  - LLM Router
  - Observability
  - Cost Tracking
  properties:
  - type: Portal
    url: https://www.langdb.ai/
  - type: GitHubOrganization
    url: https://github.com/langdb
  x-deployment:
  - cloud
  - self-host
- aid: ai-gateway:envoy-ai-gateway
  name: Envoy AI Gateway
  description: Envoy AI Gateway is an open-source extension to Envoy Proxy and Envoy Gateway, providing a Kubernetes-native
    AI traffic plane for routing, governing, and observing LLM calls across providers. Apache 2.0 licensed and CNCF-aligned.
  humanURL: https://aigateway.envoyproxy.io/
  baseURL: https://aigateway.envoyproxy.io
  tags:
  - AI Gateway
  - Envoy
  - Kubernetes
  - CNCF
  - Open Source
  properties:
  - type: Portal
    url: https://aigateway.envoyproxy.io/
  - type: Documentation
    url: https://aigateway.envoyproxy.io/docs/
  - type: GitHubRepository
    url: https://github.com/envoyproxy/ai-gateway
  - type: GitHubOrganization
    url: https://github.com/envoyproxy
  x-deployment:
  - self-host
  - opensource
  x-license: Apache-2.0
- aid: ai-gateway:gentrace
  name: Gentrace
  description: Gentrace was an AI evaluation and observability product; the company has shut down and its codebase is now
    MIT-licensed open source on GitHub. Included here for historical completeness in the AI gateway-adjacent observability
    category.
  humanURL: https://github.com/gentrace/gentrace
  tags:
  - AI Observability
  - Open Source
  - Archived
  - Evaluation
  properties:
  - type: GitHubRepository
    url: https://github.com/gentrace/gentrace
  x-deployment:
  - opensource
  x-license: MIT
  x-status: archived
- aid: ai-gateway:ai-gateway-analytics-api
  name: AI Gateway Analytics API
  description: The Analytics API from AI Gateway — 2 operation(s) for analytics.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Analytics
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-analytics-api-openapi.yml
- aid: ai-gateway:ai-gateway-apikeys-api
  name: AI Gateway APIKeys API
  description: The APIKeys API from AI Gateway — 1 operation(s) for apikeys.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - APIKeys
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-apikeys-api-openapi.yml
- aid: ai-gateway:ai-gateway-assistants-api
  name: AI Gateway Assistants API
  description: The Assistants API from AI Gateway — 1 operation(s) for assistants.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Assistants
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-assistants-api-openapi.yml
- aid: ai-gateway:ai-gateway-audio-api
  name: AI Gateway Audio API
  description: The Audio API from AI Gateway — 3 operation(s) for audio.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Audio
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-audio-api-openapi.yml
- aid: ai-gateway:ai-gateway-batches-api
  name: AI Gateway Batches API
  description: The Batches API from AI Gateway — 2 operation(s) for batches.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Batches
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-batches-api-openapi.yml
- aid: ai-gateway:ai-gateway-chat-api
  name: AI Gateway Chat API
  description: The Chat API from AI Gateway — 1 operation(s) for chat.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Chat
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-chat-api-openapi.yml
- aid: ai-gateway:ai-gateway-completions-api
  name: AI Gateway Completions API
  description: The Completions API from AI Gateway — 1 operation(s) for completions.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Completions
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-completions-api-openapi.yml
- aid: ai-gateway:ai-gateway-configs-api
  name: AI Gateway Configs API
  description: The Configs API from AI Gateway — 1 operation(s) for configs.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Configs
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-configs-api-openapi.yml
- aid: ai-gateway:ai-gateway-embeddings-api
  name: AI Gateway Embeddings API
  description: The Embeddings API from AI Gateway — 2 operation(s) for embeddings.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Embeddings
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-embeddings-api-openapi.yml
- aid: ai-gateway:ai-gateway-feedback-api
  name: AI Gateway Feedback API
  description: The Feedback API from AI Gateway — 1 operation(s) for feedback.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Feedback
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-feedback-api-openapi.yml
- aid: ai-gateway:ai-gateway-files-api
  name: AI Gateway Files API
  description: The Files API from AI Gateway — 2 operation(s) for files.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Files
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-files-api-openapi.yml
- aid: ai-gateway:ai-gateway-finetuning-api
  name: AI Gateway FineTuning API
  description: The FineTuning API from AI Gateway — 2 operation(s) for finetuning.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - FineTuning
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-finetuning-api-openapi.yml
- aid: ai-gateway:ai-gateway-guardrails-api
  name: AI Gateway Guardrails API
  description: The Guardrails API from AI Gateway — 1 operation(s) for guardrails.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Guardrails
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-guardrails-api-openapi.yml
- aid: ai-gateway:ai-gateway-images-api
  name: AI Gateway Images API
  description: The Images API from AI Gateway — 1 operation(s) for images.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Images
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-images-api-openapi.yml
- aid: ai-gateway:ai-gateway-integrations-api
  name: AI Gateway Integrations API
  description: The Integrations API from AI Gateway — 2 operation(s) for integrations.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Integrations
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-integrations-api-openapi.yml
- aid: ai-gateway:ai-gateway-logs-api
  name: AI Gateway Logs API
  description: The Logs API from AI Gateway — 2 operation(s) for logs.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Logs
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-logs-api-openapi.yml
- aid: ai-gateway:ai-gateway-mcp-api
  name: AI Gateway MCP API
  description: The MCP API from AI Gateway — 2 operation(s) for mcp.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - MCP
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-mcp-api-openapi.yml
- aid: ai-gateway:ai-gateway-policies-api
  name: AI Gateway Policies API
  description: The Policies API from AI Gateway — 2 operation(s) for policies.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Policies
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-policies-api-openapi.yml
- aid: ai-gateway:ai-gateway-prompts-api
  name: AI Gateway Prompts API
  description: The Prompts API from AI Gateway — 6 operation(s) for prompts.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Prompts
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-prompts-api-openapi.yml
- aid: ai-gateway:ai-gateway-responses-api
  name: AI Gateway Responses API
  description: The Responses API from AI Gateway — 1 operation(s) for responses.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Responses
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-responses-api-openapi.yml
- aid: ai-gateway:ai-gateway-threads-api
  name: AI Gateway Threads API
  description: The Threads API from AI Gateway — 3 operation(s) for threads.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Threads
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-threads-api-openapi.yml
- aid: ai-gateway:ai-gateway-virtualkeys-api
  name: AI Gateway VirtualKeys API
  description: The VirtualKeys API from AI Gateway — 1 operation(s) for virtualkeys.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - VirtualKeys
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-virtualkeys-api-openapi.yml
- aid: ai-gateway:ai-gateway-workspaces-api
  name: AI Gateway Workspaces API
  description: The Workspaces API from AI Gateway — 4 operation(s) for workspaces.
  humanURL: https://portkey.ai/
  baseURL: https://api.portkey.ai
  tags:
  - Workspaces
  properties:
  - type: OpenAPI
    url: openapi/ai-gateway-workspaces-api-openapi.yml
common:
- type: AgenticAccess
  url: agentic-access/ai-gateway-agentic-access.yml
- type: DomainSecurity
  url: security/ai-gateway-domain-security.yml
- type: Authentication
  url: authentication/ai-gateway-authentication.yml
- type: JSONSchema
  url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/json-schema/ai-gateway-route-schema.json
  title: AI Gateway Route Schema
- type: JSONSchema
  url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/json-schema/ai-gateway-provider-schema.json
  title: AI Gateway Provider Schema
- type: JSONSchema
  url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/json-schema/ai-gateway-policy-schema.json
  title: AI Gateway Policy Schema
- type: JSONStructure
  url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/json-structure/ai-gateway-route-structure.json
  title: AI Gateway Route Structure
- type: JSONStructure
  url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/json-structure/ai-gateway-provider-structure.json
  title: AI Gateway Provider Structure
- type: JSONStructure
  url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/json-structure/ai-gateway-policy-structure.json
  title: AI Gateway Policy Structure
- type: JSONLD
  url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/json-ld/ai-gateway-context.jsonld
- type: Vocabulary
  url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/vocabulary/ai-gateway-vocabulary.yml
- type: Examples
  url: https://raw.githubusercontent.com/api-evangelist/ai-gateway/refs/heads/main/examples/
- type: Features
  data:
  - name: Provider Abstraction
    description: A unified, typically OpenAI-compatible API surface that lets clients call any supported LLM provider without
      provider-specific SDK juggling.
  - name: Model Routing
    description: Route requests to the right model and provider based on alias, header, request content, identity, time-of-day,
      cost, or latency.
  - name: Fallback and Failover
    description: Automatically retry failed requests against backup providers or models when a primary upstream is degraded,
      rate-limited, or down.
  - name: Load Balancing and Fanout
    description: Distribute traffic across multiple providers or replicas using weighted, priority-based, or RPM/TPM-aware
      load balancing.
  - name: Response Caching
    description: Exact-match and semantic caching of model responses to cut latency and provider spend; some gateways claim
      40-70 percent cost savings.
  - name: Cost Controls and Budgets
    description: Per-user, per-team, per-key, per-project budgets, spend tracking, and hard or soft caps on token consumption.
  - name: Rate Limiting and Quotas
    description: RPM, TPM, concurrency, and per-key quotas enforced at the gateway, decoupled from each upstream provider's
      limits.
  - name: Guardrails and Prompt Firewall
    description: Prompt injection detection, jailbreak filtering, content moderation, PII redaction, and topic control applied
      to requests and responses.
  - name: Observability
    description: Request, response, token, cost, latency, error, and trace data exported via OpenTelemetry, Langfuse, Phoenix,
      Langsmith, or built-in dashboards.
  - name: Authentication and RBAC
    description: Virtual keys, JWT, OAuth2, SSO, and role-based access control over which clients can use which models with
      which budgets.
  - name: BYOK and Secret Management
    description: Bring-your-own provider API keys, with the gateway holding and injecting them so clients never see upstream
      credentials.
  - name: Multi-Tenant Governance
    description: Per-tenant isolation of keys, budgets, logs, and policies for platform teams serving multiple internal product
      teams.
  - name: MCP Federation
    description: Some AI gateways also front Model Context Protocol servers, aggregating tools and exposing a single MCP endpoint
      to agents.
- type: UseCases
  data:
  - name: Provider-Agnostic LLM Access
    description: Front many LLM providers behind one API so application teams can switch models without changing client code.
  - name: Cost Containment for AI
    description: Apply caching, routing to cheaper models, and per-team budgets to keep generative-AI spend predictable.
  - name: Reliability and Failover
    description: Survive single-provider outages by automatically failing over to backup models when the primary degrades.
  - name: Centralized AI Governance
    description: Enforce content, PII, and policy controls in one place for every AI request leaving the organization.
  - name: Observability and FinOps
    description: Attribute cost and latency to teams, projects, and users; expose token-level metrics to FinOps and SRE.
  - name: Multi-Tenant AI Platforms
    description: Build internal AI platforms where each product team gets its own virtual keys, budgets, and logs.
- type: Integrations
  data:
  - name: OpenAI
    description: Front OpenAI's GPT, embeddings, and image models behind the gateway with virtual keys and budgets.
  - name: Anthropic
    description: Route Claude requests through the gateway for fallback, caching, and central observability.
  - name: Google Gemini and Vertex AI
    description: Proxy Google Gemini and Vertex AI calls with OpenAI-format translation where supported.
  - name: AWS Bedrock
    description: Bridge OpenAI-format clients to Bedrock-hosted Anthropic, Mistral, Cohere, Meta, and Amazon models.
  - name: Azure OpenAI
    description: Route to Azure-hosted OpenAI deployments with per-region failover and key rotation.
  - name: Ollama and vLLM
    description: Front self-hosted Ollama and vLLM inference servers for hybrid cloud and on-prem inference.
  - name: OpenTelemetry
    description: Export request, token, cost, and trace data to any OTel-compatible observability backend.
  - name: Langfuse and Phoenix
    description: Stream prompts, completions, and evaluations to Langfuse and Arize Phoenix for prompt and model analytics.
  - name: Model Context Protocol
    description: Some AI gateways federate MCP servers alongside LLM routes, exposing a unified agent endpoint.
- type: Portal
  url: https://github.com/api-evangelist/ai-gateway
- type: Blog
  url: https://apievangelist.com/category/ai-gateway/
maintainers:
- FN: Kin Lane
  email: kin@apievangelist.com