Langfuse UnstableEvaluators API

The UnstableEvaluators API from Langfuse — 2 operation(s) for unstableevaluators.

OpenAPI Specification

langfuse-unstableevaluators-api-openapi.yml Raw ↑
openapi: 3.0.1
info:
  title: langfuse AnnotationQueues UnstableEvaluators API
  version: ''
  description: '## Authentication


    Authenticate with the API using [Basic Auth](https://en.wikipedia.org/wiki/Basic_access_authentication), get API keys in the project settings:


    - username: Langfuse Public Key

    - password: Langfuse Secret Key


    ## Exports


    - OpenAPI spec: https://cloud.langfuse.com/generated/api/openapi.yml'
tags:
- name: UnstableEvaluators
paths:
  /api/public/unstable/evaluators:
    post:
      description: 'Create an evaluator in the authenticated project.


        Use evaluators to define **how** Langfuse should score data: the prompt, the expected structured output, and the optional model configuration.


        Naming behavior:

        - If this is a new evaluator name in your project, Langfuse creates version `1`.

        - If the name already exists in your project, Langfuse creates the next version and returns it.

        - When a new project version is created, existing evaluation rules in that project automatically move to the newest version for that evaluator name.


        Recommended workflow:

        1. Create the evaluator.

        2. Read the returned `variables` array.

        3. Read the returned `outputDefinition.dataType` so the client knows whether future scores will be numeric, boolean, or categorical.

        4. Create one or more evaluation rules that reference the returned evaluator family using `name` and `scope`.


        Recovery guidance:

        - `422` with `code=evaluator_preflight_failed`: the evaluator cannot run with the resolved model configuration. Add a valid explicit `modelConfig`, or configure the project''s default evaluation model, then retry the same request.

        - `400` with `code=invalid_body`: the request shape is malformed. Use the structured `details.issues` array to fix the specific fields and retry.

        - `400` with `code=invalid_body` on `outputDefinition`: send `dataType`, `reasoning.description`, and `score.description`. Do not send `version`; it is not part of the public request shape.


        Unstable API note:

        - This surface may evolve while the underlying evaluation data model is being redesigned.'
      operationId: unstable_evaluators_create
      tags:
      - UnstableEvaluators
      parameters: []
      responses:
        '200':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstableEvaluator'
              examples:
                CreateEvaluatorVersion:
                  value:
                    id: evaltmpl_123
                    name: answer-correctness
                    version: 2
                    scope: project
                    type: llm_as_judge
                    prompt: 'You are grading an answer.


                      Input:

                      {{input}}


                      Output:

                      {{output}}


                      Return a score between 0 and 1.

                      '
                    variables:
                    - input
                    - output
                    outputDefinition:
                      dataType: NUMERIC
                      reasoning:
                        description: Explain why the score was assigned.
                      score:
                        description: Correctness score between 0 and 1.
                    modelConfig:
                      provider: openai
                      model: gpt-4.1-mini
                    evaluationRuleCount: 0
                    createdAt: '2026-03-30T09:00:00.000Z'
                    updatedAt: '2026-03-30T09:00:00.000Z'
        '400':
          description: ''
          content:
            application/json:
              schema: {}
        '401':
          description: ''
          content:
            application/json:
              schema: {}
        '403':
          description: ''
          content:
            application/json:
              schema: {}
        '404':
          description: ''
          content:
            application/json:
              schema: {}
        '405':
          description: ''
          content:
            application/json:
              schema: {}
        '409':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstablePublicApiError'
        '422':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstablePublicApiError'
        '429':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstablePublicApiError'
        '500':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstablePublicApiError'
      security:
      - BasicAuth: []
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/unstableCreateEvaluatorRequest'
            examples:
              CreateEvaluatorVersion:
                value:
                  name: answer-correctness
                  prompt: 'You are grading an answer.


                    Input:

                    {{input}}


                    Output:

                    {{output}}


                    Return a score between 0 and 1.

                    '
                  outputDefinition:
                    dataType: NUMERIC
                    reasoning:
                      description: Explain why the score was assigned.
                    score:
                      description: Correctness score between 0 and 1.
                  modelConfig:
                    provider: openai
                    model: gpt-4.1-mini
    get:
      description: 'List the evaluators available to the authenticated project.


        Important behavior:

        - This endpoint returns the latest version of each available evaluator.

        - Results can include evaluators from your project and Langfuse-managed evaluators.

        - If the same evaluator name exists in both places, both are returned as separate items with different `scope` values.'
      operationId: unstable_evaluators_list
      tags:
      - UnstableEvaluators
      parameters:
      - name: page
        in: query
        description: 1-based page number. Defaults to `1`.
        required: false
        schema:
          type: integer
          nullable: true
      - name: limit
        in: query
        description: Maximum number of items per page. Defaults to `50`.
        required: false
        schema:
          type: integer
          nullable: true
      responses:
        '200':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstableEvaluators'
        '400':
          description: ''
          content:
            application/json:
              schema: {}
        '401':
          description: ''
          content:
            application/json:
              schema: {}
        '403':
          description: ''
          content:
            application/json:
              schema: {}
        '404':
          description: ''
          content:
            application/json:
              schema: {}
        '405':
          description: ''
          content:
            application/json:
              schema: {}
        '429':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstablePublicApiError'
        '500':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstablePublicApiError'
      security:
      - BasicAuth: []
  /api/public/unstable/evaluators/{evaluatorId}:
    get:
      description: 'Get one evaluator by `id`.


        Use this endpoint when you want the prompt, output definition, model configuration, and derived variables for the evaluator you plan to use in an evaluation rule.'
      operationId: unstable_evaluators_get
      tags:
      - UnstableEvaluators
      parameters:
      - name: evaluatorId
        in: path
        description: Evaluator identifier returned by the evaluator endpoints.
        required: true
        schema:
          type: string
      responses:
        '200':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstableEvaluator'
        '400':
          description: ''
          content:
            application/json:
              schema: {}
        '401':
          description: ''
          content:
            application/json:
              schema: {}
        '403':
          description: ''
          content:
            application/json:
              schema: {}
        '404':
          description: ''
          content:
            application/json:
              schema: {}
        '405':
          description: ''
          content:
            application/json:
              schema: {}
        '429':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstablePublicApiError'
        '500':
          description: ''
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/unstablePublicApiError'
      security:
      - BasicAuth: []
components:
  schemas:
    utilsMetaResponse:
      title: utilsMetaResponse
      type: object
      properties:
        page:
          type: integer
          description: current page number
        limit:
          type: integer
          description: number of items per page
        totalItems:
          type: integer
          description: number of total items given the current filters/selection (if any)
        totalPages:
          type: integer
          description: number of total pages given the current limit
      required:
      - page
      - limit
      - totalItems
      - totalPages
    unstableEvaluatorModelConfig:
      title: unstableEvaluatorModelConfig
      type: object
      description: 'Optional explicit model configuration for an evaluator.


        If omitted, Langfuse uses the project''s default evaluation model.

        If provided, the model must be available to the project when the evaluator or evaluation rule is enabled.


        To discover valid configured `provider` values for a project, call `GET /api/public/llm-connections` and read the `provider` field from the returned connections.

        Use a `provider` value that matches one of the connections already configured in the same project.


        Recovery guidance:

        - If evaluator creation returns `422` with `code=evaluator_preflight_failed`, either provide a valid explicit `modelConfig` here or configure the project''s default evaluation model, then retry the same request.'
      properties:
        provider:
          type: string
          description: 'Provider identifier to use for this evaluator, for example `openai` or `anthropic`.


            To discover valid values for the current project, call `GET /api/public/llm-connections` and use one of the returned `provider` values.'
          example: openai
        model:
          type: string
          description: Model identifier exposed by the provider, for example `gpt-4.1-mini`.
          example: gpt-4.1-mini
      required:
      - provider
      - model
    unstablePublicCategoricalEvaluatorOutputDefinition:
      title: unstablePublicCategoricalEvaluatorOutputDefinition
      type: object
      properties:
        dataType:
          $ref: '#/components/schemas/unstableEvaluatorOutputDataType'
          description: Always `CATEGORICAL`.
        reasoning:
          $ref: '#/components/schemas/unstableEvaluatorOutputFieldDefinition'
        score:
          $ref: '#/components/schemas/unstablePublicCategoricalEvaluatorOutputScoreDefinition'
      required:
      - dataType
      - reasoning
      - score
    unstablePublicCategoricalEvaluatorOutputScoreDefinition:
      title: unstablePublicCategoricalEvaluatorOutputScoreDefinition
      type: object
      properties:
        description:
          type: string
        categories:
          type: array
          items:
            type: string
        shouldAllowMultipleMatches:
          type: boolean
      required:
      - description
      - categories
      - shouldAllowMultipleMatches
    unstablePublicBooleanEvaluatorOutputDefinition:
      title: unstablePublicBooleanEvaluatorOutputDefinition
      type: object
      properties:
        dataType:
          $ref: '#/components/schemas/unstableEvaluatorOutputDataType'
          description: Always `BOOLEAN`.
        reasoning:
          $ref: '#/components/schemas/unstableEvaluatorOutputFieldDefinition'
        score:
          $ref: '#/components/schemas/unstableEvaluatorOutputFieldDefinition'
      required:
      - dataType
      - reasoning
      - score
    unstableEvaluatorOutputDataType:
      title: unstableEvaluatorOutputDataType
      type: string
      enum:
      - NUMERIC
      - BOOLEAN
      - CATEGORICAL
      description: 'Structured score type returned by an evaluator.


        This controls the type of score value Langfuse stores for evaluation results:

        - `NUMERIC`: a numeric score such as `0.82`

        - `BOOLEAN`: a boolean score such as `true`

        - `CATEGORICAL`: one or more category labels from a fixed list'
    unstableCreateEvaluatorRequest:
      title: unstableCreateEvaluatorRequest
      type: object
      description: 'Request body for creating an evaluator.


        If the same `name` already exists in your project, Langfuse creates the next version and returns it.

        Existing evaluation rules in the same project are then moved to that new latest version automatically.'
      properties:
        name:
          type: string
          description: Evaluator name within the authenticated project.
          example: answer-correctness
        prompt:
          type: string
          description: Prompt template used by the evaluator.
          example: 'You are grading an answer.


            Input:

            {{input}}


            Output:

            {{output}}


            Return a score between 0 and 1.

            '
        outputDefinition:
          $ref: '#/components/schemas/unstableEvaluatorOutputDefinition'
          description: 'Structured output schema the evaluator must return.


            Always send `dataType`.

            Do not send `version`; it is an internal storage detail and not part of the public request contract.'
        modelConfig:
          $ref: '#/components/schemas/unstableEvaluatorModelConfig'
          nullable: true
          description: Optional explicit model configuration. Omit or set to `null` to use the project default evaluation model.
      required:
      - name
      - prompt
      - outputDefinition
    unstableEvaluator:
      title: unstableEvaluator
      type: object
      description: 'One evaluator that can be used for scoring.


        An evaluator describes **how** to score data:

        - prompt

        - extracted prompt variables

        - output schema

        - optional explicit model configuration


        It does not define **which** live objects are evaluated. That is the job of `evaluation-rules`.


        For agent clients, the most important fields are:

        - `variables`: use these exact names when building the evaluation-rule `mapping` array

        - `outputDefinition`: tells you the expected score type and the evaluator''s response instructions

        - `modelConfig`: tells you whether the evaluator uses the project default model (`null`) or an explicit provider/model


        Versioning behavior:

        - `GET /evaluators` returns the latest version of each available evaluator.

        - `GET /evaluators/{id}` can return an older version.

        - Evaluation rules always run against the latest version for the selected evaluator name within the same source (`project` or `managed`).'
      properties:
        id:
          type: string
          description: Identifier of this evaluator.
          example: evaltmpl_123
        name:
          type: string
          description: Evaluator name.
          example: answer-correctness
        version:
          type: integer
          description: Version number of this evaluator.
          example: 2
        scope:
          $ref: '#/components/schemas/unstableEvaluatorScope'
          description: 'Where this evaluator comes from: your project or Langfuse-managed defaults.'
        type:
          $ref: '#/components/schemas/unstableEvaluatorType'
          description: Evaluator engine type. Currently always `llm_as_judge`.
        prompt:
          type: string
          description: Prompt template used during evaluation.
          example: 'You are grading an answer.


            Input:

            {{input}}


            Output:

            {{output}}


            Return a score between 0 and 1.

            '
        variables:
          type: array
          items:
            type: string
          description: 'Variables extracted from the evaluator prompt.


            Every variable in this list must be mapped exactly once when creating an evaluation rule.'
          example:
          - input
          - output
        outputDefinition:
          $ref: '#/components/schemas/unstablePublicEvaluatorOutputDefinition'
          description: 'Structured output schema returned by this evaluator.


            Responses always include `dataType` and omit the internal output-definition `version`.

            Use `dataType` to decide how future scores should be interpreted.'
        modelConfig:
          $ref: '#/components/schemas/unstableEvaluatorModelConfig'
          nullable: true
          description: Explicit model configuration, or `null` when the project default evaluation model is used.
        evaluationRuleCount:
          type: integer
          description: Number of evaluation rules in the project that currently use this evaluator version.
          example: 0
        createdAt:
          type: string
          format: date-time
          description: Timestamp when this evaluator was created.
          example: '2026-03-30T09:00:00.000Z'
        updatedAt:
          type: string
          format: date-time
          description: Timestamp when this evaluator was last updated.
          example: '2026-03-30T09:00:00.000Z'
      required:
      - id
      - name
      - version
      - scope
      - type
      - prompt
      - variables
      - outputDefinition
      - evaluationRuleCount
      - createdAt
      - updatedAt
    unstablePublicApiErrorCode:
      title: unstablePublicApiErrorCode
      type: string
      enum:
      - authentication_failed
      - access_denied
      - invalid_request
      - invalid_query
      - invalid_body
      - invalid_filter_value
      - invalid_json_path
      - invalid_variable_mapping
      - missing_variable_mapping
      - duplicate_variable_mapping
      - resource_not_found
      - name_conflict
      - evaluator_preflight_failed
      - conflict
      - unprocessable_content
      - rate_limited
      - method_not_allowed
      - internal_error
      description: 'Machine-readable error code returned by the unstable evaluators API.


        SDKs, CLIs, and agents should branch on `code` rather than parsing the human-readable `message`.

        The HTTP status still indicates the broad error class, while `code` gives the specific failure reason.'
    unstablePublicApiError:
      title: unstablePublicApiError
      type: object
      description: 'Standard error envelope for the unstable evaluators API.


        Response handling guidance:

        - Use the HTTP status code for the broad class of failure.

        - Use `code` for precise branching in SDKs, CLIs, or agents.

        - Inspect `details` for field-level validation context such as invalid filter values, malformed JSONPath expressions, or missing variable mappings.

        - Retry only after fixing the specific issue described by `code` and `details`.'
      properties:
        message:
          type: string
          description: Human-readable description of the failure.
          example: 'Filter column "type" contains unsupported value(s): INVALID'
        code:
          $ref: '#/components/schemas/unstablePublicApiErrorCode'
          description: Stable machine-readable error code.
        details:
          $ref: '#/components/schemas/unstablePublicApiErrorDetails'
          nullable: true
          description: Optional structured error context. Inspect the populated fields based on `code`.
      required:
      - message
      - code
    unstablePublicApiErrorDetails:
      title: unstablePublicApiErrorDetails
      type: object
      description: 'Optional structured context attached to an unstable-evals error.


        The populated fields depend on the error `code`:

        - request parsing failures populate `issues`

        - filter validation failures populate `field`, `column`, `invalidValues`, and `allowedValues`

        - variable mapping failures populate `field`, `variable`, or `variables`

        - JSONPath validation failures populate `field`, `variable`, and `value`

        - evaluator preflight failures populate `evaluatorName`, `provider`, and `model`

        - rate limiting populates `retryAfterSeconds`, `limit`, `remaining`, and `resetAt`'
      properties:
        issues:
          type: array
          items:
            $ref: '#/components/schemas/unstablePublicApiValidationIssue'
          nullable: true
          description: Validation issues for malformed request bodies or query parameters.
        field:
          type: string
          nullable: true
          description: Path-like reference to the failing field, for example `mapping[1].jsonPath`.
        column:
          type: string
          nullable: true
          description: Filter column that failed validation.
        invalidValues:
          type: array
          items:
            type: string
          nullable: true
          description: Unsupported values supplied by the caller.
        allowedValues:
          type: array
          items:
            type: string
          nullable: true
          description: Allowed values for the failing filter column.
        variable:
          type: string
          nullable: true
          description: Evaluator variable involved in the failure.
        variables:
          type: array
          items:
            type: string
          nullable: true
          description: Multiple evaluator variables involved in the failure, for example missing mappings.
        value:
          type: string
          nullable: true
          description: Raw invalid value supplied by the caller.
        evaluatorName:
          type: string
          nullable: true
          description: Evaluator name used during preflight validation.
        provider:
          type: string
          nullable: true
          description: Provider resolved during evaluator preflight, if any.
        model:
          type: string
          nullable: true
          description: Model resolved during evaluator preflight, if any.
        retryAfterSeconds:
          type: integer
          nullable: true
          description: Suggested retry delay for rate-limited requests.
        limit:
          type: integer
          nullable: true
          description: Numeric limit associated with the failure, for example the active evaluation-rule cap or the current rate-limit window.
        remaining:
          type: integer
          nullable: true
          description: Remaining requests in the current rate-limit window.
        resetAt:
          type: string
          nullable: true
          description: ISO-8601 timestamp when the current rate-limit window resets.
    unstablePublicNumericEvaluatorOutputDefinition:
      title: unstablePublicNumericEvaluatorOutputDefinition
      type: object
      properties:
        dataType:
          $ref: '#/components/schemas/unstableEvaluatorOutputDataType'
          description: Always `NUMERIC`.
        reasoning:
          $ref: '#/components/schemas/unstableEvaluatorOutputFieldDefinition'
        score:
          $ref: '#/components/schemas/unstableEvaluatorOutputFieldDefinition'
      required:
      - dataType
      - reasoning
      - score
    unstablePublicApiValidationIssue:
      title: unstablePublicApiValidationIssue
      type: object
      description: 'One validation issue returned for malformed request bodies or query parameters.


        This mirrors the most important parts of a Zod issue: a machine-readable `code`,

        a human-readable `message`, and a structured `path`.'
      properties:
        code:
          type: string
          description: Machine-readable validation issue code emitted by the server validator.
        message:
          type: string
          description: Human-readable explanation of the validation failure.
        path:
          type: array
          items: {}
          description: Path to the invalid field, for example `["mapping", 0, "jsonPath"]`.
      required:
      - code
      - message
      - path
    unstableEvaluators:
      title: unstableEvaluators
      type: object
      properties:
        data:
          type: array
          items:
            $ref: '#/components/schemas/unstableEvaluator'
        meta:
          $ref: '#/components/schemas/utilsMetaResponse'
      required:
      - data
      - meta
    unstableEvaluatorScope:
      title: unstableEvaluatorScope
      type: string
      enum:
      - project
      - managed
      description: 'Where an evaluator comes from.


        - `project`: created in your project

        - `managed`: provided by Langfuse'
    unstableEvaluatorType:
      title: unstableEvaluatorType
      type: string
      enum:
      - llm_as_judge
      description: 'The evaluator engine type.


        The unstable public API currently supports only LLM-as-a-judge evaluators.'
    unstableEvaluatorOutputDefinition:
      title: unstableEvaluatorOutputDefinition
      oneOf:
      - type: object
        allOf:
        - type: object
          properties:
            dataType:
              type: string
              enum:
              - NUMERIC
        - $ref: '#/components/schemas/unstablePublicNumericEvaluatorOutputDefinition'
        required:
        - dataType
      - type: object
        allOf:
        - type: object
          properties:
            dataType:
              type: string
              enum:
              - BOOLEAN
        - $ref: '#/components/schemas/unstablePublicBooleanEvaluatorOutputDefinition'
        required:
        - dataType
      - type: object
        allOf:
        - type: object
          properties:
            dataType:
              type: string
              enum:
              - CATEGORICAL
        - $ref: '#/components/schemas/unstablePublicCategoricalEvaluatorOutputDefinition'
        required:
        - dataType
      description: 'Structured output definition to send when creating an evaluator.


        Agent guidance:

        - `dataType` is required.

        - Do not send `version`; that is an internal storage detail and is not part of the public request contract.

        - For `NUMERIC` and `BOOLEAN`, provide `reasoning.description` and `score.description`.

        - For `CATEGORICAL`, also provide `score.categories` and `score.shouldAllowMultipleMatches`.'
    unstableEvaluatorOutputFieldDefinition:
      title: unstableEvaluatorOutputFieldDefinition
      type: object
      properties:
        description:
          type: string
          description: Human-readable instructions for what the evaluator should return in this field.
      required:
      - description
    unstablePublicEvaluatorOutputDefinition:
      title: unstablePublicEvaluatorOutputDefinition
      oneOf:
      - type: object
        allOf:
        - type: object
          properties:
            dataType:
              type: string
              enum:
              - NUMERIC
        - $ref: '#/components/schemas/unstablePublicNumericEvaluatorOutputDefinition'
        required:
        - dataType
      - type: object
        allOf:
        - type: object
          properties:
            dataType:
              type: string
              enum:
              - BOOLEAN
        - $ref: '#/components/schemas/unstablePublicBooleanEvaluatorOutputDefinition'
        required:
        - dataType
      - type: object
        allOf:
        - type: object
          properties:
            dataType:
              type: string
              enum:
              - CATEGORICAL
        - $ref: '#/components/schemas/unstablePublicCategoricalEvaluatorOutputDefinition'
        required:
        - dataType
      description: 'Evaluator output definition returned by the public API.


        This response always includes `dataType` and never includes an internal output-definition `version`.

        Legacy stored evaluator definitions are normalized into this shape before they are returned.


        Use this response shape when deciding how to interpret future evaluation scores:

        - `NUMERIC`: expect numeric score values

        - `BOOLEAN`: expect `true` / `false`

        - `CATEGORICAL`: expect one or more values from `score.categories`'
  securitySchemes:
    BasicAuth:
      type: http
      scheme: basic