Ollama Chat Completions API

Generate chat completions using the OpenAI-compatible chat endpoint with multi-turn conversation support.

Business capability
Artificial Intelligence Management BC-610.60

Operations 2

POST /chat/completions Ollama Create chat completion #
POST /api/chat Ollama Generate a chat completion #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/ollama-chat-completions-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

ollama-chat-completions-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: Ollama OpenAI Compatibility Chat Completions API
  description: Ollama provides compatibility with parts of the OpenAI API, allowing existing applications built for OpenAI to connect to locally-running models through Ollama. Supported endpoints include chat completions, completions, embeddings, models, images, and the Responses API.
  version: 0.1.0
  contact:
    name: Ollama Team
    url: https://ollama.com
  license:
    name: MIT
    url: https://opensource.org/licenses/MIT
servers:
- url: http://localhost:11434/v1
  description: Local Ollama Server (OpenAI-compatible)
security:
- bearerAuth: []
tags:
- name: Chat Completions
  description: Generate chat completions using the OpenAI-compatible chat endpoint with multi-turn conversation support.
paths:
  /chat/completions:
    post:
      operationId: createChatCompletion
      summary: Ollama Create chat completion
      description: Creates a model response for the given chat conversation. Supports streaming, JSON mode, structured output, vision, and tool calling. Compatible with the OpenAI Chat Completions API format.
      tags:
      - Chat Completions
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ChatCompletionRequest'
      responses:
        '200':
          description: Successful chat completion response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ChatCompletionResponse'
            text/event-stream:
              schema:
                $ref: '#/components/schemas/ChatCompletionStreamResponse'
        '400':
          description: Bad Request
        '404':
          description: Model not found
  /api/chat:
    post:
      operationId: generateChatCompletion
      summary: Ollama Generate a chat completion
      description: Generate the next message in a chat with a provided model. This is a streaming endpoint that returns a sequence of JSON objects by default. Supports multi-turn conversations, tool calling, vision input, and structured output.
      tags:
      - Chat Completions
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ChatRequest'
      responses:
        '200':
          description: Successful chat response. Returns a stream of JSON objects when streaming is enabled, or a single JSON object when disabled.
          content:
            application/x-ndjson:
              schema:
                $ref: '#/components/schemas/ChatResponse'
            application/json:
              schema:
                $ref: '#/components/schemas/ChatResponse'
        '400':
          description: Bad Request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
        '404':
          description: Model not found
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
components:
  schemas:
    ContentPart:
      type: object
      description: A content part within a multimodal message.
      required:
      - type
      properties:
        type:
          type: string
          description: The type of content part.
          enum:
          - text
          - image_url
        text:
          type: string
          description: The text content when type is text.
        image_url:
          type: object
          description: The image URL or base64 data when type is image_url.
          properties:
            url:
              type: string
              description: URL of the image or a base64-encoded data URI.
    ChatCompletionStreamResponse:
      type: object
      description: A streaming chat completion chunk.
      properties:
        id:
          type: string
          description: A unique identifier for the chat completion.
        object:
          type: string
          description: The object type, always chat.completion.chunk.
          const: chat.completion.chunk
        created:
          type: integer
          description: Unix timestamp of when the chunk was created.
        model:
          type: string
          description: The model used for the completion.
        choices:
          type: array
          description: A list of chat completion chunk choices.
          items:
            type: object
            properties:
              index:
                type: integer
                description: The index of the choice.
              delta:
                type: object
                description: The delta content for this chunk.
                properties:
                  role:
                    type: string
                    description: The role of the author.
                  content:
                    type: string
                    description: The content delta.
                  tool_calls:
                    type: array
                    description: Tool call deltas.
                    items:
                      $ref: '#/components/schemas/OpenAIToolCall'
              finish_reason:
                type:
                - string
                - 'null'
                description: The finish reason, if applicable.
    UsageStats:
      type: object
      description: Token usage statistics for the request.
      properties:
        prompt_tokens:
          type: integer
          description: Number of tokens in the prompt.
        completion_tokens:
          type: integer
          description: Number of tokens in the generated completion.
        total_tokens:
          type: integer
          description: Total number of tokens used in the request.
    ChatCompletionRequest:
      type: object
      description: Request body for creating a chat completion in OpenAI-compatible format.
      required:
      - model
      - messages
      properties:
        model:
          type: string
          description: The model to use for chat completion.
        messages:
          type: array
          description: A list of messages comprising the conversation so far.
          items:
            $ref: '#/components/schemas/ChatCompletionMessage'
        temperature:
          type: number
          description: Sampling temperature between 0 and 2. Higher values make output more random.
          minimum: 0.0
          maximum: 2.0
        top_p:
          type: number
          description: Nucleus sampling parameter. Considers tokens with top_p probability mass.
          minimum: 0.0
          maximum: 1.0
        max_tokens:
          type: integer
          description: Maximum number of tokens to generate in the response.
        frequency_penalty:
          type: number
          description: Penalty for token frequency to reduce repetition.
          minimum: -2.0
          maximum: 2.0
        presence_penalty:
          type: number
          description: Penalty for token presence to encourage topic diversity.
          minimum: -2.0
          maximum: 2.0
        seed:
          type: integer
          description: Random seed for deterministic generation.
        stop:
          description: Sequences where the API will stop generating further tokens.
          oneOf:
          - type: string
          - type: array
            items:
              type: string
        stream:
          type: boolean
          description: If true, partial message deltas are sent as server-sent events.
          default: false
        stream_options:
          type: object
          description: Options for streaming responses.
          properties:
            include_usage:
              type: boolean
              description: If true, includes usage information in the stream.
        response_format:
          type: object
          description: Specifies the format of the response. Use type json_object for JSON mode or json_schema for structured output.
          properties:
            type:
              type: string
              description: The response format type.
              enum:
              - text
              - json_object
              - json_schema
            json_schema:
              type: object
              description: The JSON Schema for structured output.
              additionalProperties: true
        tools:
          type: array
          description: A list of tools the model may call.
          items:
            $ref: '#/components/schemas/OpenAIToolDefinition'
    ChatCompletionResponse:
      type: object
      description: Response object from a chat completion request.
      properties:
        id:
          type: string
          description: A unique identifier for the chat completion.
        object:
          type: string
          description: The object type, always chat.completion.
          const: chat.completion
        created:
          type: integer
          description: Unix timestamp of when the completion was created.
        model:
          type: string
          description: The model used for the completion.
        choices:
          type: array
          description: A list of chat completion choices.
          items:
            $ref: '#/components/schemas/ChatCompletionChoice'
        usage:
          $ref: '#/components/schemas/UsageStats'
    ChatCompletionMessage:
      type: object
      description: A message in the chat conversation.
      required:
      - role
      properties:
        role:
          type: string
          description: The role of the message author.
          enum:
          - system
          - user
          - assistant
          - tool
        content:
          description: The content of the message. Can be a string or an array of content parts for multimodal input.
          oneOf:
          - type: string
          - type: array
            items:
              $ref: '#/components/schemas/ContentPart'
        name:
          type: string
          description: An optional name for the participant.
        tool_calls:
          type: array
          description: Tool calls generated by the model.
          items:
            $ref: '#/components/schemas/OpenAIToolCall'
        tool_call_id:
          type: string
          description: Tool call that this message is responding to.
    OpenAIToolDefinition:
      type: object
      description: A tool definition in OpenAI-compatible format.
      required:
      - type
      - function
      properties:
        type:
          type: string
          description: The type of tool. Currently only function is supported.
          enum:
          - function
        function:
          type: object
          description: The function definition.
          required:
          - name
          properties:
            name:
              type: string
              description: The name of the function.
            description:
              type: string
              description: A description of what the function does.
            parameters:
              type: object
              description: The function parameters as a JSON Schema object.
              additionalProperties: true
    OpenAIToolCall:
      type: object
      description: A tool call generated by the model.
      properties:
        id:
          type: string
          description: A unique identifier for the tool call.
        type:
          type: string
          description: The type of tool call.
          enum:
          - function
        function:
          type: object
          description: The function call details.
          properties:
            name:
              type: string
              description: The name of the function to call.
            arguments:
              type: string
              description: The arguments to pass to the function as a JSON string.
    ChatCompletionChoice:
      type: object
      description: A single chat completion choice.
      properties:
        index:
          type: integer
          description: The index of the choice in the list.
        message:
          $ref: '#/components/schemas/ChatCompletionMessage'
        finish_reason:
          type: string
          description: The reason the model stopped generating tokens.
          enum:
          - stop
          - length
          - tool_calls
    Logprob:
      type: object
      description: Log probability information for a generated token.
      properties:
        token:
          type: string
          description: The token text.
        logprob:
          type: number
          description: The log probability of the token.
        bytes:
          type: array
          description: The byte representation of the token.
          items:
            type: integer
        top_logprobs:
          type: array
          description: The most likely alternative tokens and their log probabilities at this position.
          items:
            type: object
            properties:
              token:
                type: string
                description: The alternative token text.
              logprob:
                type: number
                description: The log probability of the alternative token.
              bytes:
                type: array
                description: The byte representation of the alternative token.
                items:
                  type: integer
    ChatRequest:
      type: object
      description: Request body for generating a chat completion from a conversation.
      required:
      - model
      - messages
      properties:
        model:
          type: string
          description: The name of the model to use for chat completion.
        messages:
          type: array
          description: The messages of the chat conversation. Each message has a role and content, with optional images and tool calls.
          items:
            $ref: '#/components/schemas/ChatMessage'
        tools:
          type: array
          description: A list of tool definitions that the model may call during the conversation for function calling.
          items:
            $ref: '#/components/schemas/ToolDefinition'
        format:
          description: The format to return a response in. Accepts the string json for simple JSON mode, or a JSON Schema object for structured output.
          oneOf:
          - type: string
            enum:
            - json
          - type: object
        stream:
          type: boolean
          description: When true, returns a stream of partial responses as newline-delimited JSON. When false, returns a single response object.
          default: true
        think:
          description: When true, returns separate thinking output in addition to the generated content. Can also be set to a string value to control thinking verbosity.
          oneOf:
          - type: boolean
          - type: string
            enum:
            - high
            - medium
            - low
        keep_alive:
          description: How long the model stays loaded in memory after the request.
          oneOf:
          - type: string
          - type: number
        options:
          $ref: '#/components/schemas/ModelOptions'
        logprobs:
          type: boolean
          description: Whether to return log probabilities of the output tokens.
        top_logprobs:
          type: integer
          description: Number of most likely tokens to return at each token position.
          minimum: 0
    ChatResponse:
      type: object
      description: Response object from a chat completion request. When streaming, partial objects are returned until done is true.
      properties:
        model:
          type: string
          description: The model used to generate the chat response.
        created_at:
          type: string
          format: date-time
          description: The ISO 8601 timestamp when the response was created.
        message:
          $ref: '#/components/schemas/ChatMessage'
        done:
          type: boolean
          description: Indicates whether the chat response has finished.
        done_reason:
          type: string
          description: The reason the generation stopped.
        total_duration:
          type: integer
          description: Total time spent generating the response in nanoseconds.
        load_duration:
          type: integer
          description: Time spent loading the model in nanoseconds.
        prompt_eval_count:
          type: integer
          description: Number of tokens in the prompt that were evaluated.
        prompt_eval_duration:
          type: integer
          description: Time spent evaluating the prompt in nanoseconds.
        eval_count:
          type: integer
          description: Number of tokens generated in the response.
        eval_duration:
          type: integer
          description: Time spent generating output tokens in nanoseconds.
        logprobs:
          type: array
          description: Log probability information for generated tokens.
          items:
            $ref: '#/components/schemas/Logprob'
    ToolCall:
      type: object
      description: A tool call requested by the model during function calling.
      properties:
        function:
          type: object
          description: The function to call with its name and arguments.
          properties:
            name:
              type: string
              description: The name of the function to call.
            arguments:
              type: object
              description: The arguments to pass to the function as key-value pairs.
              additionalProperties: true
    ToolDefinition:
      type: object
      description: A definition of a tool that the model may call during conversation for function calling.
      required:
      - type
      - function
      properties:
        type:
          type: string
          description: The type of tool. Currently only function is supported.
          enum:
          - function
        function:
          type: object
          description: The function definition including name, description, and parameter schema.
          required:
          - name
          properties:
            name:
              type: string
              description: The name of the function.
            description:
              type: string
              description: A description of what the function does.
            parameters:
              type: object
              description: A JSON Schema object defining the function parameters.
              additionalProperties: true
    ChatMessage:
      type: object
      description: A message in a chat conversation, containing a role, content, and optional images or tool calls.
      required:
      - role
      - content
      properties:
        role:
          type: string
          description: The role of the message sender.
          enum:
          - system
          - user
          - assistant
          - tool
        content:
          type: string
          description: The text content of the message.
        thinking:
          type: string
          description: The thinking content when extended thinking is enabled.
        images:
          type: array
          description: A list of base64-encoded images attached to the message for multimodal models.
          items:
            type: string
            format: byte
        tool_calls:
          type: array
          description: Tool calls requested by the assistant for function calling.
          items:
            $ref: '#/components/schemas/ToolCall'
    ErrorResponse:
      type: object
      description: An error response returned when a request fails.
      properties:
        error:
          type: string
          description: A human-readable error message describing the problem.
    ModelOptions:
      type: object
      description: Runtime options that control text generation behavior. These override the model's default parameter values.
      properties:
        seed:
          type: integer
          description: Random seed for reproducible generation. Set to a specific value for deterministic output.
        temperature:
          type: number
          description: Controls the creativity of responses. Higher values produce more varied output. Range 0.0 to 2.0.
          minimum: 0.0
          maximum: 2.0
        top_k:
          type: integer
          description: Limits token selection to the top K most probable tokens. Lower values produce more focused output.
          minimum: 0
        top_p:
          type: number
          description: Nucleus sampling probability cutoff. Limits selection to the smallest set of tokens whose cumulative probability exceeds this value.
          minimum: 0.0
          maximum: 1.0
        min_p:
          type: number
          description: Minimum probability threshold relative to the most likely token. Tokens below this threshold are filtered out.
          minimum: 0.0
          maximum: 1.0
        stop:
          type: array
          description: A list of stop sequences. Generation halts when any of these strings are produced.
          items:
            type: string
        num_ctx:
          type: integer
          description: The maximum context window size in tokens.
        num_predict:
          type: integer
          description: The maximum number of tokens to generate in the response.
        repeat_penalty:
          type: number
          description: Penalty applied to repeated tokens to reduce repetition.
        repeat_last_n:
          type: integer
          description: Number of recent tokens to consider for repeat penalty.
        tfs_z:
          type: number
          description: Tail-free sampling parameter. Higher values reduce the impact of less probable tokens.
        mirostat:
          type: integer
          description: Enable Mirostat sampling for perplexity control. 0 is disabled, 1 uses Mirostat, 2 uses Mirostat 2.0.
          enum:
          - 0
          - 1
          - 2
        mirostat_tau:
          type: number
          description: Target entropy for Mirostat sampling.
        mirostat_eta:
          type: number
          description: Learning rate for Mirostat sampling.
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: API key authentication. The key is accepted but not validated by Ollama. Use any value such as ollama.
externalDocs:
  description: Ollama OpenAI Compatibility Documentation
  url: https://docs.ollama.com/api/openai-compatibility