Agenta Evaluations API

Run evaluations of variants against testsets.

Operations 4

POST /evaluations/runs/query List and filter evaluation runs. #
POST /evaluations/runs/ Create an evaluation run. #
POST /evaluations/results/query Query evaluation results. #
POST /evaluations/metrics/query Query aggregated evaluation metrics. #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/agenta-evaluations-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

agenta-evaluations-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: Agenta Applications Evaluations API
  description: Agenta is an open-source LLMOps platform for prompt management, LLM evaluation, and LLM observability. This specification documents the public cloud REST API surface used to manage applications and variants, fetch and deploy versioned prompt configurations, run evaluations and configure evaluators, manage testsets, and ingest and query observability traces. All endpoints are authenticated with an Agenta API key passed in the Authorization header. Agenta is MIT licensed and may also be self-hosted.
  termsOfService: https://agenta.ai/terms
  contact:
    name: Agenta Support
    url: https://agenta.ai/
    email: team@agenta.ai
  license:
    name: MIT
    url: https://github.com/Agenta-AI/agenta/blob/main/LICENSE
  version: '1.0'
servers:
- url: https://cloud.agenta.ai/api
  description: Agenta Cloud (US)
- url: https://eu.cloud.agenta.ai/api
  description: Agenta Cloud (EU)
security:
- ApiKeyAuth: []
tags:
- name: Evaluations
  description: Run evaluations of variants against testsets.
paths:
  /evaluations/runs/query:
    post:
      operationId: queryEvaluationRuns
      tags:
      - Evaluations
      summary: List and filter evaluation runs.
      requestBody:
        required: false
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/EvaluationRunQueryRequest'
      responses:
        '200':
          description: A list of evaluation runs.
          content:
            application/json:
              schema:
                type: array
                items:
                  $ref: '#/components/schemas/EvaluationRun'
        '401':
          $ref: '#/components/responses/Unauthorized'
  /evaluations/runs/:
    post:
      operationId: createEvaluationRun
      tags:
      - Evaluations
      summary: Create an evaluation run.
      description: Starts an evaluation run that executes one or more variants against a testset and scores the outputs with the configured evaluators.
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/EvaluationRunCreateRequest'
      responses:
        '200':
          description: The created evaluation run.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/EvaluationRun'
        '401':
          $ref: '#/components/responses/Unauthorized'
  /evaluations/results/query:
    post:
      operationId: queryEvaluationResults
      tags:
      - Evaluations
      summary: Query evaluation results.
      requestBody:
        required: false
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/EvaluationResultQueryRequest'
      responses:
        '200':
          description: A list of evaluation results.
          content:
            application/json:
              schema:
                type: array
                items:
                  $ref: '#/components/schemas/EvaluationResult'
        '401':
          $ref: '#/components/responses/Unauthorized'
  /evaluations/metrics/query:
    post:
      operationId: queryEvaluationMetrics
      tags:
      - Evaluations
      summary: Query aggregated evaluation metrics.
      requestBody:
        required: false
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/EvaluationMetricsQueryRequest'
      responses:
        '200':
          description: Aggregated metrics for the matching runs.
          content:
            application/json:
              schema:
                type: array
                items:
                  $ref: '#/components/schemas/EvaluationMetrics'
        '401':
          $ref: '#/components/responses/Unauthorized'
components:
  schemas:
    Error:
      type: object
      properties:
        detail:
          oneOf:
          - type: string
          - type: object
    Windowing:
      type: object
      description: Cursor / time-window pagination controls.
      properties:
        next:
          type: string
        limit:
          type: integer
          default: 100
        oldest:
          type: string
          format: date-time
        newest:
          type: string
          format: date-time
        order:
          type: string
          enum:
          - ascending
          - descending
    EvaluationMetricsQueryRequest:
      type: object
      properties:
        metrics:
          type: object
          properties:
            run_id:
              type: string
              format: uuid
    EvaluationMetrics:
      type: object
      properties:
        id:
          type: string
          format: uuid
        run_id:
          type: string
          format: uuid
        evaluator_slug:
          type: string
        value:
          type: number
        data:
          type: object
          additionalProperties: true
    EvaluationRunQueryRequest:
      type: object
      properties:
        run:
          type: object
          properties:
            application_id:
              type: string
              format: uuid
        include_archived:
          type: boolean
          default: false
        windowing:
          $ref: '#/components/schemas/Windowing'
    EvaluationRunCreateRequest:
      type: object
      required:
      - run
      properties:
        run:
          type: object
          properties:
            name:
              type: string
            testset_id:
              type: string
              format: uuid
            variant_ids:
              type: array
              items:
                type: string
                format: uuid
            evaluator_ids:
              type: array
              items:
                type: string
                format: uuid
    EvaluationResultQueryRequest:
      type: object
      properties:
        result:
          type: object
          properties:
            run_id:
              type: string
              format: uuid
        windowing:
          $ref: '#/components/schemas/Windowing'
    EvaluationResult:
      type: object
      properties:
        id:
          type: string
          format: uuid
        run_id:
          type: string
          format: uuid
        scenario_id:
          type: string
          format: uuid
        status:
          type: string
        outputs:
          type: object
          additionalProperties: true
    Lifecycle:
      type: object
      description: Audit metadata attached to most Agenta resources.
      properties:
        created_at:
          type: string
          format: date-time
        updated_at:
          type: string
          format: date-time
        created_by_id:
          type: string
          format: uuid
        updated_by_id:
          type: string
          format: uuid
    EvaluationRun:
      type: object
      properties:
        id:
          type: string
          format: uuid
        name:
          type: string
        status:
          type: string
          enum:
          - pending
          - running
          - finished
          - failed
          - cancelled
        testset_id:
          type: string
          format: uuid
        lifecycle:
          $ref: '#/components/schemas/Lifecycle'
  responses:
    Unauthorized:
      description: Missing or invalid API key.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
  securitySchemes:
    ApiKeyAuth:
      type: http
      scheme: bearer
      description: 'Agenta API key sent in the Authorization header. Generate keys from the Agenta web app under Settings > API keys. The value is passed as `Authorization: ApiKey <key>` (Bearer-style header credential).'