LocalAI Monitoring API

The monitoring API from LocalAI — 16 operation(s) for monitoring.

Operations 16

GET /api/backend-logs List models with backend logs #
GET /api/backend-logs/{modelId} Get backend logs for a model #
POST /api/backend-logs/{modelId}/clear Clear backend logs for a model #
GET /api/backend-traces List backend operation traces #
POST /api/backend-traces/clear Clear backend traces #
GET /api/backend-traces/{id} Get one backend operation trace #
GET /api/traces List API request/response traces #
POST /api/traces/clear Clear API traces #
GET /api/traces/summary Summarize recent API traces #
GET /api/traces/{id} Get one API trace #
POST /backend/load Pre-load a model into memory #
GET /backend/monitor Backend monitor endpoint #
POST /backend/shutdown Backend shutdown endpoint #
GET /metrics Prometheus metrics endpoint #
GET /system Show the LocalAI instance information #
GET /ws/backend-logs/{modelId} Stream backend logs via WebSocket #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/localai-monitoring-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

localai-monitoring-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  description: The LocalAI Rest API.
  title: LocalAI Monitoring API
  contact:
    name: LocalAI
    url: https://localai.io
  license:
    name: MIT
    url: https://raw.githubusercontent.com/mudler/LocalAI/master/LICENSE
  version: 2.0.0
servers:
- url: /
tags:
- name: Monitoring
paths:
  /api/backend-logs:
    get:
      description: Returns a sorted list of model IDs that have captured backend process output
      tags:
      - Monitoring
      summary: List models with backend logs
      responses:
        '200':
          description: Model IDs with logs
          content:
            application/json:
              schema:
                type: array
                items:
                  type: string
      operationId: getApiBackendLogs
      x-operation-id-source: derived
  /api/backend-logs/{modelId}:
    get:
      description: Returns all captured log lines (stdout/stderr) for the specified model's backend process
      tags:
      - Monitoring
      summary: Get backend logs for a model
      parameters:
      - description: Model ID
        name: modelId
        in: path
        required: true
        schema:
          type: string
      responses:
        '200':
          description: Log lines
          content:
            application/json:
              schema:
                type: array
                items:
                  $ref: '#/components/schemas/model.BackendLogLine'
      operationId: getApiBackendLogsByModelId
      x-operation-id-source: derived
  /api/backend-logs/{modelId}/clear:
    post:
      description: Removes all captured log lines for the specified model's backend process
      tags:
      - Monitoring
      summary: Clear backend logs for a model
      parameters:
      - description: Model ID
        name: modelId
        in: path
        required: true
        schema:
          type: string
      responses:
        '204':
          description: Logs cleared
      operationId: postApiBackendLogsByModelIdClear
      x-operation-id-source: derived
  /api/backend-traces:
    get:
      description: Returns a bounded, newest-first page of captured backend traces (LLM calls, embeddings, TTS, etc). The heavy body and data fields are omitted unless full=true; fetch them per-trace from /api/backend-traces/{id}. Paging metadata is returned in the X-Total-Count, X-Trace-Offset and X-Trace-Limit headers.
      tags:
      - Monitoring
      summary: List backend operation traces
      parameters:
      - description: Maximum entries to return (default 50, max 1000, 0 for all)
        name: limit
        in: query
        schema:
          type: integer
      - description: Number of entries to skip (default 0)
        name: offset
        in: query
        schema:
          type: integer
      - description: Include the body and data payloads (default false)
        name: full
        in: query
        schema:
          type: boolean
      responses:
        '200':
          description: Backend operation traces
          content:
            application/json:
              schema:
                type: object
                additionalProperties: true
      operationId: getApiBackendTraces
      x-operation-id-source: derived
  /api/backend-traces/clear:
    post:
      description: Removes all captured backend operation traces from the buffer
      tags:
      - Monitoring
      summary: Clear backend traces
      responses:
        '204':
          description: Traces cleared
      operationId: postApiBackendTracesClear
      x-operation-id-source: derived
  /api/backend-traces/{id}:
    get:
      description: Returns a single captured backend trace, including the body and data payloads omitted from the list response
      tags:
      - Monitoring
      summary: Get one backend operation trace
      parameters:
      - description: Trace ID
        name: id
        in: path
        required: true
        schema:
          type: string
      responses:
        '200':
          description: Backend operation trace
          content:
            application/json:
              schema:
                type: object
                additionalProperties: true
        '404':
          description: Trace not found
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/schema.ErrorResponse'
      operationId: getApiBackendTracesById
      x-operation-id-source: derived
  /api/traces:
    get:
      description: Returns a bounded, newest-first page of captured API exchange traces. Request and response bodies plus headers are omitted unless full=true; fetch them per-trace from /api/traces/{id}. Paging metadata is returned in the X-Total-Count, X-Trace-Offset and X-Trace-Limit headers.
      tags:
      - Monitoring
      summary: List API request/response traces
      parameters:
      - description: Maximum entries to return (default 50, max 1000, 0 for all)
        name: limit
        in: query
        schema:
          type: integer
      - description: Number of entries to skip (default 0)
        name: offset
        in: query
        schema:
          type: integer
      - description: Include request/response bodies and headers (default false)
        name: full
        in: query
        schema:
          type: boolean
      responses:
        '200':
          description: Traced API exchanges
          content:
            application/json:
              schema:
                type: object
                additionalProperties: true
      operationId: getApiTraces
      x-operation-id-source: derived
  /api/traces/clear:
    post:
      description: Removes all captured API request/response traces from the buffer
      tags:
      - Monitoring
      summary: Clear API traces
      responses:
        '204':
          description: Traces cleared
      operationId: postApiTracesClear
      x-operation-id-source: derived
  /api/traces/summary:
    get:
      description: Returns request, failure and latency totals over a recent window, plus a bucketed series for sparklines. Exists so callers wanting three numbers do not have to fetch the whole trace list and count it themselves.
      tags:
      - Monitoring
      summary: Summarize recent API traces
      parameters:
      - description: Window in hours (default 24, max 168)
        name: hours
        in: query
        schema:
          type: integer
      responses:
        '200':
          description: Counted trace totals
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/middleware.TraceSummary'
      operationId: getApiTracesSummary
      x-operation-id-source: derived
  /api/traces/{id}:
    get:
      description: Returns a single captured API exchange, including the request and response bodies omitted from the list response
      tags:
      - Monitoring
      summary: Get one API trace
      parameters:
      - description: Trace ID
        name: id
        in: path
        required: true
        schema:
          type: string
      responses:
        '200':
          description: Traced API exchange
          content:
            application/json:
              schema:
                type: object
                additionalProperties: true
        '404':
          description: Trace not found
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/schema.ErrorResponse'
      operationId: getApiTracesById
      x-operation-id-source: derived
  /backend/load:
    post:
      description: Loads the named model (or, for a realtime pipeline, all of its sub-models) into memory so subsequent requests pay no cold-start cost. The inverse of /backend/shutdown.
      tags:
      - Monitoring
      summary: Pre-load a model into memory
      responses:
        '200':
          description: Model loaded
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/schema.ModelLoadResponse'
        '400':
          description: Missing model name
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/schema.ModelLoadResponse'
        '500':
          description: Load failed (Loaded lists any sub-models that did load)
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/schema.ModelLoadResponse'
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/schema.ModelLoadRequest'
        description: Model to load
        required: true
      operationId: postBackendLoad
      x-operation-id-source: derived
  /backend/monitor:
    get:
      tags:
      - Monitoring
      summary: Backend monitor endpoint
      parameters:
      - description: Name of the model to monitor
        name: model
        in: query
        required: true
        schema:
          type: string
      responses:
        '200':
          description: Response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/proto.StatusResponse'
      operationId: getBackendMonitor
      x-operation-id-source: derived
  /backend/shutdown:
    post:
      tags:
      - Monitoring
      summary: Backend shutdown endpoint
      responses: {}
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/schema.BackendMonitorRequest'
        description: Backend statistics request
        required: true
      operationId: postBackendShutdown
      x-operation-id-source: derived
  /metrics:
    get:
      tags:
      - Monitoring
      summary: Prometheus metrics endpoint
      responses:
        '200':
          description: Prometheus metrics
          content:
            text/plain:
              schema:
                type: string
      operationId: getMetrics
      x-operation-id-source: derived
  /system:
    get:
      tags:
      - Monitoring
      summary: Show the LocalAI instance information
      responses:
        '200':
          description: Response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/schema.SystemInformationResponse'
      operationId: getSystem
      x-operation-id-source: derived
  /ws/backend-logs/{modelId}:
    get:
      description: Opens a WebSocket connection for real-time backend log streaming. Sends an initial batch of existing lines (type "initial"), then streams new lines as they appear (type "line"). Supports ping/pong keepalive.
      tags:
      - Monitoring
      summary: Stream backend logs via WebSocket
      parameters:
      - description: Model ID
        name: modelId
        in: path
        required: true
        schema:
          type: string
      responses: {}
      operationId: getWsBackendLogsByModelId
      x-operation-id-source: derived
components:
  schemas:
    proto.StatusResponse_State:
      type: integer
      format: int32
      enum:
      - 0
      - 1
      - 2
      - -1
      x-enum-varnames:
      - StatusResponse_UNINITIALIZED
      - StatusResponse_BUSY
      - StatusResponse_READY
      - StatusResponse_ERROR
    schema.ModelLoadRequest:
      type: object
      properties:
        model:
          type: string
    schema.SysInfoModel:
      type: object
      properties:
        backend:
          description: 'Backend is the engine serving this model. The loader knows only the ID,

            so it is resolved from the model''s config; empty when the model was

            loaded without one (a loose file, or a config since removed).'
          type: string
        id:
          type: string
    schema.SystemInformationResponse:
      type: object
      properties:
        backends:
          description: available backend engines
          type: array
          items:
            type: string
        loaded_models:
          description: currently loaded models
          type: array
          items:
            $ref: '#/components/schemas/schema.SysInfoModel'
    proto.MemoryUsageData:
      type: object
      properties:
        breakdown:
          type: object
          additionalProperties:
            type: integer
            format: int64
        total:
          type: integer
    schema.BackendMonitorRequest:
      type: object
      properties:
        model:
          type: string
    middleware.TraceBucket:
      type: object
      properties:
        count:
          type: integer
        errors:
          type: integer
        start:
          type: string
    model.BackendLogLine:
      type: object
      properties:
        stream:
          description: '"stdout" or "stderr"'
          type: string
        text:
          type: string
        timestamp:
          type: string
    schema.APIError:
      type: object
      properties:
        code: {}
        message:
          type: string
        param:
          type: string
        type:
          type: string
    middleware.TraceSummary:
      type: object
      properties:
        buckets:
          type: array
          items:
            $ref: '#/components/schemas/middleware.TraceBucket'
        errors:
          type: integer
        p95_ms:
          type: integer
        total:
          type: integer
        window_hours:
          type: integer
    schema.ErrorResponse:
      type: object
      properties:
        error:
          $ref: '#/components/schemas/schema.APIError'
    proto.StatusResponse:
      type: object
      properties:
        memory:
          $ref: '#/components/schemas/proto.MemoryUsageData'
        state:
          $ref: '#/components/schemas/proto.StatusResponse_State'
    schema.ModelLoadResponse:
      type: object
      properties:
        loaded:
          description: 'Loaded lists the model names actually resident in memory after the call.

            For a pipeline model these are its sub-models, not the pipeline name.'
          type: array
          items:
            type: string
        message:
          description: Message is a short human-readable status ("model loaded", or an error).
          type: string
  securitySchemes:
    BearerAuth:
      type: apiKey
      name: Authorization
      in: header