Lamini Inference API

Text completion and streaming generation endpoints.

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/lamini-inference-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no email required.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

lamini-inference-api-openapi.yml Raw ↑
openapi: 3.0.1
info:
  title: Lamini Platform Classify Inference API
  description: REST API for the Lamini enterprise LLM platform covering inference (completions), fine-tuning and Memory Tuning jobs, classification, and embeddings over open base and tuned models. All requests are authenticated with a Bearer API key and served from https://api.lamini.ai. Endpoints and request fields are derived from the official Lamini Python client (github.com/lamini-ai/lamini) and the Lamini REST API documentation.
  termsOfService: https://www.lamini.ai/terms
  contact:
    name: Lamini Support
    url: https://www.lamini.ai
  version: '1.0'
servers:
- url: https://api.lamini.ai
security:
- bearerAuth: []
tags:
- name: Inference
  description: Text completion and streaming generation endpoints.
paths:
  /v1/completions:
    post:
      operationId: createCompletion
      tags:
      - Inference
      summary: Generate a completion
      description: Generate a text completion from a base or tuned model. Supports a typed output schema via output_type for structured generation.
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/CompletionRequest'
      responses:
        '200':
          description: A generated completion.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/CompletionResponse'
        '401':
          description: Missing or invalid API key.
        '429':
          description: Rate limit exceeded.
  /v3/streaming_completions:
    post:
      operationId: createStreamingCompletion
      tags:
      - Inference
      summary: Generate a streaming completion
      description: Generate a completion as an incremental stream of token chunks for the provided prompt and model.
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/CompletionRequest'
      responses:
        '200':
          description: A stream of completion chunks.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/CompletionResponse'
components:
  schemas:
    CompletionRequest:
      type: object
      required:
      - prompt
      - model_name
      properties:
        prompt:
          oneOf:
          - type: string
          - type: array
            items:
              type: string
          description: One or more input prompts.
        model_name:
          type: string
          description: Base or tuned model identifier to generate from.
        output_type:
          type: object
          additionalProperties: true
          description: Optional typed output schema for structured generation.
        max_tokens:
          type: integer
          nullable: true
        max_new_tokens:
          type: integer
          nullable: true
        cache_id:
          type: string
          nullable: true
    CompletionResponse:
      type: object
      properties:
        output:
          oneOf:
          - type: string
          - type: object
            additionalProperties: true
          description: Generated text or structured output.
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: 'Lamini platform API key passed as Authorization: Bearer <API_KEY>. Requests may also include a Lamini-Version header.'