SambaNova Systems Chat completions API

The Chat completions API from SambaNova Systems — 1 operation(s) for chat completions.

Operations 1

POST /chat/completions Create chat-based completion #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/sambanova-systems-chat-completions-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no email required.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

sambanova-systems-chat-completions-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: SambaNova cloud Chat completions API
  description: SambaNova cloud API Specification
  version: 1.2.0
  termsOfService: https://sambanova.ai/cloud-end-user-license-agreement
  contact:
    email: info@sambanova.ai
    name: SambaNova information
  license:
    name: Apache 2.0
    url: https://www.apache.org/licenses/LICENSE-2.0.html
servers:
- url: https://api.sambanova.ai/v1
security:
- api_key: []
tags:
- name: Chat completions
paths:
  /chat/completions:
    post:
      operationId: createChatCompletion
      tags:
      - Chat completions
      summary: Create chat-based completion
      security:
      - api_key: []
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ChatCompletionRequest'
        description: Chat prompt and parameters
        required: true
      responses:
        '200':
          description: 'Successful Response, Returns a ChatCompletionResponse object (non-streaming), or a stream of server-sent ChatCompletionStreamResponse object events ending with a response.completed event (when stream: true).'
          content:
            application/json:
              schema:
                oneOf:
                - $ref: '#/components/schemas/ChatCompletionResponse'
                - $ref: '#/components/schemas/ChatCompletionStreamResponse'
        '400':
          description: Bad Request - Missing or invalid parameters
          content:
            text/plain:
              schema:
                type: string
                example: 'Invalid request body:  - missing property ''model''"'
            application/json:
              schema:
                oneOf:
                - $ref: '#/components/schemas/GeneralError'
                - $ref: '#/components/schemas/SimpleError'
              examples:
                empty_messages:
                  summary: Empty messages array
                  value:
                    error:
                      message: 'Invalid ''messages'': empty array. Expected an array with minimum length 1, but got an empty array instead.'
                      type: invalid_request_error
                      param: messages
                      code: empty_array
                    request_id: d79vdot7os633dverihg
                invalid_role:
                  summary: Invalid message role value
                  value:
                    error:
                      message: 'Invalid value: ''developerrr''. Supported values are: ''system'', ''assistant'', ''user'', ''function'', ''tool'', and ''developer''.'
                      type: invalid_request_error
                      param: messages[0].role
                      code: invalid_value
                    request_id: d79vdit7os6aemenqav0
                null_content:
                  summary: Null user message content
                  value:
                    error:
                      message: 'Invalid value for ''content'': expected a string, got null.'
                      type: invalid_request_error
                      param: messages.[5].content
                      code: null
                    request_id: d79vcsd7os6aemenqa80
                missing_model:
                  summary: No model parameter provided
                  value:
                    error:
                      message: you must provide a model parameter
                      type: invalid_request_error
                      param: null
                      code: null
                    request_id: a4fa849e
                invalid_json:
                  summary: Malformed JSON body
                  value:
                    error:
                      code: null
                      message: 'We could not parse the JSON body of your request. (HINT: This likely means you aren''t using your HTTP library correctly. A JSON payload is expected, but what was sent was not valid JSON.)'
                      param: null
                      type: invalid_request_error
                    request_id: 3f7db127
                simple:
                  summary: Simple error (unhandled cases)
                  value:
                    error: Unhandled error
        '401':
          description: Unauthorized access
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/GeneralError'
              examples:
                invalid_api_key:
                  summary: Invalid API key provided
                  value:
                    error:
                      message: 'Incorrect API key provided: *****. You can find your API key at https://cloud.sambanova.ai/apis.'
                      type: invalid_request_error
                      param: null
                      code: invalid_api_key
                    request_id: d79vcq57os633dverhi0
                missing_api_key:
                  summary: No API key provided
                  value:
                    error:
                      message: 'You didn''t provide an API key. You need to provide your API key in an Authorization header using Bearer auth (i.e. Authorization: Bearer YOUR_KEY).'
                      type: invalid_request_error
                      param: null
                      code: null
                    request_id: d7a0dct7os633dvesi60
        '404':
          description: Not found - model does not exist or wrong endpoint called
          content:
            application/json:
              schema:
                oneOf:
                - $ref: '#/components/schemas/GeneralError'
                - $ref: '#/components/schemas/SimpleError'
              examples:
                model_not_found:
                  summary: Model does not exist or is not accessible
                  value:
                    error:
                      code: model_not_found
                      message: The model `abc` does not exist or you do not have access to it.
                      param: model
                      type: invalid_request_error
                    request_id: 887a8227
                simple:
                  summary: Simple error (unhandled cases e.g. wrong endpoint)
                  value:
                    error: Not found
        '408':
          description: Request timeout
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/SimpleError'
        '410':
          description: Gone - model is no longer available (deprecated or removed)
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/SimpleError'
        '429':
          description: Too Many Requests
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/GeneralError'
        '500':
          description: Internal Server Error. Unexpected issue on server side.
          content:
            text/plain:
              schema:
                type: string
        '503':
          description: Service Temporarily Unavailable
          content:
            text/plain:
              schema:
                type: string
                example: Service Temporarily Unavailable
      x-codeSamples:
      - lang: JavaScript
        source: "import SambaNova from 'sambanova';\n\nconst client = new SambaNova({\n  apiKey: process.env['SAMBANOVA_API_KEY'], // This is the default and can be omitted\n});\n\nconst completion = await client.chat.completions.create({\n  messages: [{ content: 'create a poem using palindromes', role: 'user' }],\n  model: 'gpt-oss-120b',\n});\n\nconsole.log(completion);"
      - lang: Python
        source: "import os\nfrom sambanova import SambaNova\n\nclient = SambaNova(\n    api_key=os.environ.get(\"SAMBANOVA_API_KEY\"),  # This is the default and can be omitted\n)\nfor completion in client.chat.completions.create(\n    messages=[{\n        \"content\": \"create a poem using palindromes\",\n        \"role\": \"user\",\n    }],\n    model=\"gpt-oss-120b\",\n):\n  print(completion)"
components:
  schemas:
    ImageContent:
      title: Image Content
      type: object
      additionalProperties: true
      properties:
        type:
          title: Type
          type: string
          description: type of content to send. in this case `image_url`.
          enum:
          - image_url
          const: image_url
        image_url:
          title: Image Url
          type: object
          properties:
            url:
              title: Url
              type: string
              description: Either a URL of the image or the base64 encoded image data.  currently only base64 encoded image supported
      required:
      - type
      - image_url
    ChatCompletionResponse:
      title: Chat Completion Response
      type: object
      description: chat completion response returned by the model
      properties:
        choices:
          title: Choices
          type: array
          items:
            $ref: '#/components/schemas/ChatCompletionChoice'
          minItems: 1
        created:
          title: Created
          type: number
          description: The Unix timestamp (in seconds) of when the chat completion was created.
        id:
          title: Id
          type: string
          description: A unique identifier for the chat completion.
        model:
          title: Model
          description: The model used for the chat completion.
          type: string
        object:
          title: Object
          type: string
          description: The object type, always `chat.completion`.
          enum:
          - chat.completion
          const: chat.completion
        system_fingerprint:
          title: System fingerprint
          type: string
          description: Backend configuration that the model runs with.
        usage:
          $ref: '#/components/schemas/Usage'
      required:
      - choices
      - created
      - id
      - model
      - object
      - system_fingerprint
      - usage
      examples:
      - choices:
        - finish_reason: stop
          index: 0
          message:
            content: 'In madam moon''s silver glow,  Aha, a palindrome to know,  Radar spins, a circular tale,  Level heads prevail, without fail.  A man, a plan, a canal, Panama!  Able was I ere I saw Elba,  A Santa at NASA, a curious sight,  Do geese see God, in the pale moonlight?  Mr. Owl ate my metal worm,  Do nine men interpret? Nine men, I nod,  Never odd or even, a palindrome''s might,  Madam, in Eden, I''m Adam.  Aibohphobia, a fear to confess,  A palindrome''s symmetry, I must address,  Refer, a word that reads the same,  A palindrome''s beauty, in its circular game.  In the stillness of the night,  Ava, a palindrome, shining bright,  Hannah, a name that reads the same,  A palindrome''s magic, in its circular flame.  Note: Please keep in mind that creating a poem using palindromes can be a challenging task,  and the resulting poem may not be as cohesive or flowing as one that doesn''t rely on palindromes.  However, I hope you enjoy the attempt!'
            role: assistant
        created: 1737583288.6076705
        id: 83a7809d-e18f-44f9-9ab7-2bc494c6c661
        model: gpt-oss-120b
        object: chat.completion
        system_fingerprint: fastcoe
        usage:
          acceptance_rate: 4.058139324188232
          completion_tokens: 350
          completion_tokens_after_first_per_sec: 248.09314856382406
          completion_tokens_after_first_per_sec_first_ten: 249.67922929952655
          completion_tokens_per_sec: 238.91966176995348
          end_time: 1737583289.7345645
          is_last_response: true
          prompt_tokens: 43
          start_time: 1737583288.264706
          time_to_first_token: 0.06312894821166992
          total_latency: 1.4649275719174653
          total_tokens: 393
          total_tokens_per_sec: 268.27264878740493
      - choices:
        - finish_reason: stop
          index: 0
          message:
            content: Hello! How can I assist you today?
            role: assistant
          logprobs:
            content:
            - token: ' Hello'
              logprob: -0.00012340000000000002
              bytes:
              - 32
              - 72
              - 101
              - 108
              - 108
              - 111
              top_logprobs:
              - token: ' Hello'
                logprob: -0.00012340000000000002
                bytes:
                - 32
                - 72
                - 101
                - 108
                - 108
                - 111
              - token: ' Hi'
                logprob: -8.243
                bytes:
                - 32
                - 72
                - 105
            - token: '!'
              logprob: -0.0023456
              bytes:
              - 33
              top_logprobs:
              - token: '!'
                logprob: -0.0023456
                bytes:
                - 33
              - token: ','
                logprob: -6.712
                bytes:
                - 44
        created: 1737583290.1234567
        id: 94b8910e-f29a-55a0-0bc8-3cd505d7d772
        model: gpt-oss-120b
        object: chat.completion
        system_fingerprint: fastcoe
        usage:
          acceptance_rate: 3.5
          completion_tokens: 9
          completion_tokens_after_first_per_sec: 245.3
          completion_tokens_per_sec: 230.1
          end_time: 1737583290.2345679
          is_last_response: true
          prompt_tokens: 15
          start_time: 1737583290.1234567
          time_to_first_token: 0.058
          total_latency: 0.111
          total_tokens: 24
          total_tokens_per_sec: 216.2
    FunctionObject:
      title: Function Object
      additionalProperties: true
      type: object
      properties:
        name:
          title: Name
          type: string
          description: The name of the function to be called. Must be a-z, A-Z, 0-9, or contain underscores and dashes.
        description:
          title: Description
          type: string
          description: A description of what the function does, used by the model to choose when and how to call the function.
          nullable: true
        parameters:
          $ref: '#/components/schemas/FunctionParameters'
      required:
      - name
    ChatCompletionRequest:
      title: Chat Completion Request
      type: object
      description: chat completions request object
      additionalProperties: true
      properties:
        model:
          title: Model
          description: The model ID to use (e.g. gpt-oss-120b).  See available [models](https://docs.sambanova.ai/docs/en/models/sambacloud-models)
          anyOf:
          - type: string
          - enum:
            - Meta-Llama-3.3-70B-Instruct
            - Meta-Llama-3.2-1B-Instruct
            - Meta-Llama-3.2-3B-Instruct
            - Llama-3.2-11B-Vision-Instruct
            - Llama-3.2-90B-Vision-Instruct
            - Meta-Llama-3.1-8B-Instruct
            - Meta-Llama-3.1-70B-Instruct
            - Meta-Llama-3.1-405B-Instruct
            - Qwen2.5-Coder-32B-Instruct
            - Qwen2.5-72B-Instruct
            - QwQ-32B-Preview
            - Meta-Llama-Guard-3-8B
            - DeepSeek-R1
            - DeepSeek-R1-0528
            - DeepSeek-V3-0324
            - DeepSeek-V3.1
            - DeepSeek-V3.1-cb
            - DeepSeek-V3.1-Terminus
            - DeepSeek-V3.2
            - DeepSeek-R1-Distill-Llama-70B
            - Llama-4-Maverick-17B-128E-Instruct
            - Llama-4-Scout-17B-16E-Instruct
            - Qwen3-32B
            - Qwen3-235B
            - Llama-3.3-Swallow-70B-Instruct-v0.4
            - gpt-oss-120b
            - ALLaM-7B-Instruct-preview
            - MiniMax-M2.5
            - MiniMax-M2.7
            - gemma-3-12b-it
        messages:
          title: Messages
          type: array
          description: A list of messages comprising the conversation so far.
          items:
            anyOf:
            - $ref: '#/components/schemas/SystemMessage'
            - $ref: '#/components/schemas/UserMessage'
            - $ref: '#/components/schemas/AssistantMessage'
            - $ref: '#/components/schemas/ToolMessage'
          minItems: 1
          examples:
          - - role: user
              content: create a poem using palindromes
        max_tokens:
          title: Max Tokens
          type: integer
          description: The maximum number of tokens that can be generated in the chat completion.  The total length of input tokens and generated tokens is limited by  the model's context length.
          nullable: true
          example: 2048
        max_completion_tokens:
          title: Max Completion Tokens
          type: integer
          description: The maximum number of tokens that can be generated in the chat completion.  The total length of input tokens and generated tokens is limited by  the model's context length.
          nullable: true
          example: 2048
        temperature:
          title: Temperature
          type: number
          description: What sampling temperature to use, determines the degree of randomness in the response.  between 0 and 2, Higher values like 0.8 will make the output more random,  while lower values like 0.2 will make it more focused and deterministic.  Is recommended altering this, top_p or top_k but not more than one of these.
          minimum: 0
          maximum: 2
          default: 0.7
          nullable: true
          example: 0.7
        top_p:
          title: Top P
          type: number
          description: Cumulative probability for token choices. An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. Is recommended altering this, top_k or temperature but  not more than one of these.
          minimum: 0
          maximum: 1
          nullable: true
          example: 1
        top_k:
          title: Top K
          type: integer
          description: Amount limit of token choices. An alternative to sampling with temperature, the model considers the results of the first K tokens with higher probability.  So 10 means only the first 10 tokens with higher probability are considered. Is recommended altering this, top_p or temperature but not more than one of these.
          minimum: 1
          maximum: 100
          nullable: true
          example: 5
        presence_penalty:
          title: Presence Penalty
          type: number
          description: Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics. Not currently implemented; accepted for API compatibility
          maximum: 2
          minimum: -2
          default: 0
          nullable: true
        frequency_penalty:
          title: Frequency Penalty
          type: number
          description: Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim. Not currently implemented; accepted for API compatibility
          maximum: 2
          minimum: -2
          default: 0
        do_sample:
          title: do_sample
          type: boolean
          description: If true, sampling is enabled during output generation. If false, deterministic decoding is used.
          nullable: true
        stop:
          title: Stop
          description: Sequences where the API will stop generating tokens. The returned text will not contain the stop sequence.
          oneOf:
          - type: string
            example: '

              '
            nullable: true
          - type: array
            items:
              type: string
              example: '["\n"]'
          nullable: true
        stream:
          title: Stream
          type: boolean
          description: 'If set, partial message deltas will be sent. Tokens will be sent as data-only  [server-sent events](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events#Event_stream_format) as they become available, with the stream terminated by a `data: [DONE]` message.'
          default: false
          nullable: true
        stream_options:
          title: StreamOptions
          type: object
          description: Options for streaming response. Only set this when setting stream as true
          additionalProperties: true
          properties:
            include_usage:
              anyOf:
              - type: boolean
              title: Include Usage
              description: Whether to include the usage metrics in a final chunk or not
              nullable: true
          nullable: true
        response_format:
          title: Response Format
          description: 'An object specifying the format that the model must output.   Setting to `{ "type": "json_object"}` enables JSON mode,    which will check the message the model generates is valid JSON.  **Important:** when using JSON mode, you **must** also instruct the model to produce JSON yourself  via a system or user message. Setting to `{ "type": "json_schema", "json_schema": {<your_schema>}"}` enables JSON schema mode,    which will check the message the model generates is valid object of type <your_schema>.  Setting to `{ "type": "text"}` is equivalent to the default plain text generation'
          oneOf:
          - $ref: '#/components/schemas/ResponseFormatJSONSchema'
          - $ref: '#/components/schemas/ResponseFormatJSONObject'
          - $ref: '#/components/schemas/ResponseFormatText'
          discriminator:
            propertyName: type
            mapping:
              json_schema: '#/components/schemas/ResponseFormatJSONSchema'
              json_object: '#/components/schemas/ResponseFormatJSONObject'
              text: '#/components/schemas/ResponseFormatText'
          nullable: true
        reasoning_effort:
          title: Reasoning Effort
          type: string
          description: Value specifying the amount of reasoning the model is allowed to do, increasing it will increase the number of output reasoning tokens generated by the model, but will improve quality of the responses. allowed values are 'low', 'medium', 'high'
          enum:
          - low
          - medium
          - high
          nullable: true
        tool_choice:
          title: Tool Choice
          description: 'Controls which (if any) tool is called by the model.  `none` means the model will not call any tool and instead generates a message.  `auto` means the model can pick between generating a message or calling one or more tools.  `required` means the model must call one or more tools.  Specifying a particular tool via `{"type": "function", "function": {"name": "my_function"}}` forces  the model to call that tool.'
          anyOf:
          - enum:
            - none
            - auto
            - required
            type: string
          - $ref: '#/components/schemas/ToolChoiceObject'
          nullable: true
        parallel_tool_calls:
          title: Parallel Tool Calls
          type: boolean
          description: Whether to enable parallel function calling during tool use, This is not yet supported by our models.
          nullable: true
        tools:
          title: tools
          type: array
          description: A list of tools the model may call.  Use this to provide a list of functions the model may generate JSON inputs for.
          items:
            $ref: '#/components/schemas/Tool'
          maxItems: 128
          nullable: true
        chat_template_kwargs:
          title: Chat Template Kwargs
          type: object
          description: A dictionary of additional keyword arguments to pass into the chat template.  Use this to provide extra context or parameters that the model's chat template  can process. Keys must be strings; values may be any valid JSON type.
          properties:
            enable_thinking:
              $ref: '#/components/schemas/EnableThinking'
          additionalProperties: true
          nullable: true
          example:
            enable_thinking: true
        logprobs:
          title: Logprobs
          type: boolean
          description: Whether to return log probabilities of the output tokens or not. If true, returns the log  probabilities of each output token returned in the `content` of `message`.
          default: false
          nullable: true
        top_logprobs:
          title: Top Logprobs
          type: integer
          description: An integer between 0 and 20 specifying the number of most likely tokens to return at each token  position, each with an associated log probability. `logprobs` must be set to `true` if this parameter is used.
          maximum: 20
          minimum: 0
          nullable: true
        n:
          title: N
          type: integer
          description: 'How many completions to generate for each prompt.

            **Note:** Because this parameter generates many completions, it can quickly consume your token quota. Use carefully and ensure that you have reasonable settings for `max_tokens`.'
          maximum: 8
          minimum: 1
          default: 1
          nullable: true
          example: 1
        logit_bias:
          title: Logit Bias
          type: object
          description: Modify the likelihood of specified tokens appearing in the generation. Accepts a JSON object that maps tokens (specified by their  token ID in the model tokenizer) to an associated bias value from -100 to 100. Mathematically, the bias is added to the logits generated by the model prior to sampling.  Values between -1 and 1 should decrease or increase likelihood of selection; values like -100 or 100 should result in a ban or exclusive selection of the relevant token.
          nullable: true
        seed:
          title: Seed
          description: 'If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same `seed` and parameters should return the same result.

            Determinism is not guaranteed, and you should refer to the `system_fingerprint` response parameter to monitor changes in the backend.'
          nullable: true
          type: integer
      required:
      - model
      - messages
      examples:
      - messages:
        - role: user
          content: create a poem using palindromes
        model: gpt-oss-120b
      - messages:
        - role: system
          content: You are a helpful assistant developed by SambaNova systems
        - role: user
          content: create a poem using palindromes
        max_tokens: 2048
        model: gpt-oss-120b
        stream: false
        temperature: 0.7
        top_p: 1
    ChatCompletionStreamResponse:
      title: Chat Completion Stream Response
      type: object
      description: streamed chunk of a chat completion response returned by the model
      additionalProperties: true
      properties:
        choices:
          title: Choices
          type: array
          description: A list of chat completion choices.
          items:
            $ref: '#/components/schemas/ChatCompletionChunkChoice'
          minItems: 0
          nullable: true
        created:
          title: Created
          type: number
          description: The Unix timestamp (in seconds) of when the chat completion was created.
        id:
          title: Id
          type: string
          description: A unique identifier for the chat completion.
        model:
          title: Model
          description: The model used for the chat completion.
          type: string
        object:
          title: Object
          type: string
          description: The object type, always `chat.completion.chunk`.
          enum:
          - chat.completion.chunk
          const: chat.completion.chunk
        system_fingerprint:
          title: System fingerprint
          type: string
          description: Backend configuration that the model runs with.
        usage:
          $ref: '#/components/schemas/Usage'
      required:
      - choices
      - created
      - id
      - model
      - object
      - system_fingerprint
      examples:
      - choices:
        - delta:
            role: assistant
            content: ''
          index: 0
          finish_reason: null
          logprobs: null
        created: 1737642515.6076705
        id: 15fcaba4-1c4a-48fc-bc6e-05b55ddfb30b
        model: gpt-oss-120b
        object: chat.completion.chunk
        system_fingerprint: fastcoe
      - choices:
        - delta:
            role: assistant
            content: in
          index: 0
          finish_reason: null
          logprobs: null
        created: 1737642515.6076705
        id: 15fcaba4-1c4a-48fc-bc6e-05b55ddfb30b
        model: gpt-oss-120b
        object: chat.completion.chunk
        system_fingerprint: fastcoe
      - choices:
        - delta:
            content: ''
          index: 0
          finish_reason: length
          logprobs: null
        created: 1737642515.6076705
        id: 15fcaba4-1c4a-48fc-bc6e-05b55ddfb30b
        model: gpt-oss-120b
        object: chat.completion.chunk
        system_fingerprint: fastcoe
      - choices: []
        created: 1737642515.6076705
        id: 15fcaba4-1c4a-48fc-bc6e-05b55ddfb30b
        model: gpt-oss-120b
        object: chat.completion.chunk
        system_fingerprint: fastcoe
        usage:
          acceptance_rate: 3
          completion_tokens: 100
          completion_tokens_after_first_per_sec: 262.2771255106759
          completion_tokens_after_first_per_sec_first_ten: 266.98193514986144
          completion_tokens_after_first_per_sec_graph: 266.98193514986144
          completion_tokens_per_sec: 217.87260449707574
          end_time: 1737642515.9077535
          is_last_response: true
          prompt_tokens: 43
          prompt_tokens_details:
            cached_tokens: 0
          start_time: 1737642515.4458635
          time_to_first_token: 0.0844266414642334
          total_latency: 0.4589838186899821
          total_tokens: 143
          total_tokens_per_sec: 311.55782443081836
      - choices:
        - delta:
            content: ' Hello'
          index: 0
          finish_reason: null
          logprobs:
            content:
            - token: ' Hello'
              logprob: -0.00012340000000000002
              bytes:
              - 32
              - 72
              - 101
              - 108
              - 108
              - 111
              top_logprobs:
              - token: ' Hello'
                logprob: -0.00012340000000000002
                bytes:
                - 32
                - 72
                - 101
                - 108
                - 108
                - 111
              - token: ' Hi'
                logprob: -8.243
                bytes:
                - 32
                - 72
                - 105
        created: 1737642516.1076705
        id: 15fcaba4-1c4a-48fc-bc6e-05b55ddfb30c
        model: gpt-oss-120b
        object: chat.completion.chunk
        system_fingerprint: fastcoe
    ResponseFormatJSONSchema:
      title: ResponseFormatJSONSchema
      type: object
      additionalProperties: true
      description: Specifies that the model should produce output conforming to a given JSON schema.
      properties:
        json_schema:
          $ref: '#/components/schemas/JSONSchema'
        type:
          const: json_schema
          enum:
          - json_schema
          title: Type
          type: string
      required:
      - json_schema
      - type
      example:
        type: json_schema
        json_schema:
          name: User
          description: JSON schema for a simple user object
          strict: false
          schema:
            type: object
            properties:
              id:
                type: string
                description: Unique identifier for the user
              name:
                type: string
                description: Full name of the user
            required:
            - id
            - name
    TopLogProbs:
      title: TopLogProbs
      type: object
      additionalProperties: true
      properties:
        bytes:
          title: Bytes
          anyOf:
          - items:
              type: integer
            type: array
          - type: 'n

# --- truncated at 32 KB (55 KB total) ---
# Full source: https://raw.githubusercontent.com/api-evangelist/sambanova-systems/refs/heads/main/openapi/sambanova-systems-chat-completions-api-openapi.yml