Snowflake cortex-inference API

The cortex-inference API from Snowflake — 2 operation(s) for cortex-inference.

Operations 2

GET /api/v2/cortex/models Returns the Llms Available for the Current Session #
POST /api/v2/cortex/inference:complete Perform Llm Text Completion Inference. #

Documentation

📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/account
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/alert
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/api-integration
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/catalog-integration
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/compute-pool
📖
Documentation
https://docs.snowflake.com/en/user-guide/snowflake-cortex/cortex-analyst/rest-api
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/cortex-inference
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/cortex-search-service
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/database-role
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/database
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/dynamic-table
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/event-table
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/external-volume
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/function
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/grant
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/iceberg-table
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/image-repository
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/managed-account
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/network-policy
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/notebook
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/notification-integration
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/pipe
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/procedure
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/result
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/role
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/schema
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/service
📖
Documentation
https://docs.snowflake.com/en/developer-guide/sql-api/index
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/stage
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/stream
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/table
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/task
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/user-defined-function
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/user
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/view
📖
Documentation
https://docs.snowflake.com/en/developer-guide/snowflake-rest-api/reference/warehouse

Specifications

Schemas & Data

Other Resources

🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/account-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/alert-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/api-integration-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/catalog-integration-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/compute-pool-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/cortex-analyst-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/cortex-inference-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/cortex-search-service-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/database-role-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/database-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/dynamic-table-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/event-table-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/external-volume-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/function-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/grant-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/iceberg-table-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/image-repository-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/managed-account-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/network-policy-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/notebook-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/notification-integration-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/pipe-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/procedure-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/role-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/schema-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/service-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/sqlapi-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/snowflake-sql-rest-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/stage-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/stream-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/table-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/task-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/user-defined-function-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/user-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/view-context.jsonld
🔗
JSONLD
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/json-ld/warehouse-context.jsonld
🔗
Overlay
https://raw.githubusercontent.com/api-evangelist/snowflake/refs/heads/main/overlays/snowflake-cortex-inference-overlay.yaml

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/snowflake-cortex-inference-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

snowflake-cortex-inference-api-openapi.yml Raw ↑
openapi: 3.2.0
info:
  title: Snowflake Cortex Inference API
  version: 0.1.0
  contact:
    name: Snowflake, Inc.
    url: https://snowflake.com
    email: support@snowflake.com
  description: 'Operations tagged cortex-inference across 2 of this provider''s published API definitions: cortex-inference.yaml, cortex-inference.yaml. Each path carries the servers of the definition it was published in.'
tags:
- name: cortex-inference
paths:
  /api/v2/cortex/models:
    get:
      summary: Returns the Llms Available for the Current Session
      tags:
      - cortex-inference
      description: Returns the LLMs available for the current session
      operationId: getModels
      requestBody:
        required: false
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/GetModelsRequest'
            examples:
              GetmodelsRequestExample:
                summary: Default getModels request
                x-microcks-default: true
                value:
                  models:
                  - example_value
      responses:
        '200':
          description: OK
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/GetModelsResponse'
              examples:
                Getmodels200Example:
                  summary: Default getModels 200 response
                  x-microcks-default: true
                  value:
                    models:
                    - example_value
        '400':
          $ref: common.yaml#/components/responses/400BadRequest
        '401':
          $ref: common.yaml#/components/responses/401Unauthorized
        '403':
          $ref: common.yaml#/components/responses/403Forbidden
        '404':
          $ref: common.yaml#/components/responses/404NotFound
        '405':
          $ref: common.yaml#/components/responses/405MethodNotAllowed
        '500':
          $ref: common.yaml#/components/responses/500InternalServerError
        '503':
          $ref: common.yaml#/components/responses/503ServiceUnavailable
        '504':
          $ref: common.yaml#/components/responses/504GatewayTimeout
      x-microcks-operation:
        delay: 0
        dispatcher: FALLBACK
      security:
      - KeyPair: []
      - ExternalOAuth: []
      - SnowflakeOAuth: []
  /api/v2/cortex/inference:complete:
    post:
      summary: Perform Llm Text Completion Inference.
      tags:
      - cortex-inference
      description: Perform LLM text completion inference, similar to snowflake.cortex.Complete.
      operationId: cortexLLMInferenceComplete
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/CompleteRequest'
            examples:
              CortexllminferencecompleteRequestExample:
                summary: Default cortexLLMInferenceComplete request
                x-microcks-default: true
                value:
                  model: example_value
                  messages:
                  - role: example_value
                    content: example_value
                    content_list:
                    - {}
                  temperature: 42.5
                  top_p: 42.5
                  max_tokens: 10
                  max_output_tokens: 10
                  response_format:
                    type: json
                    schema: example_value
                  guardrails:
                    enabled: true
                    response_when_unsafe: example_value
                  tools:
                  - {}
                  tool_choice: {}
                  provisioned_throughput_id: '500123'
                  sf-ml-xp-inflight-prompt-action: example_value
                  sf-ml-xp-inflight-prompt-client-id: '500123'
                  sf-ml-xp-inflight-prompt-public-key: example_value
                  stream: true
      responses:
        '200':
          description: OK
          content:
            text/event-stream:
              schema:
                $ref: '#/components/schemas/StreamingCompleteResponse'
              examples:
                Cortexllminferencecomplete200Example:
                  summary: Default cortexLLMInferenceComplete 200 response
                  x-microcks-default: true
                  value: {}
        '400':
          $ref: common.yaml#/components/responses/400BadRequest
        '401':
          $ref: common.yaml#/components/responses/401Unauthorized
        '403':
          $ref: common.yaml#/components/responses/403Forbidden
        '404':
          $ref: common.yaml#/components/responses/404NotFound
        '405':
          $ref: common.yaml#/components/responses/405MethodNotAllowed
        '500':
          $ref: common.yaml#/components/responses/500InternalServerError
        '503':
          $ref: common.yaml#/components/responses/503ServiceUnavailable
        '504':
          $ref: common.yaml#/components/responses/504GatewayTimeout
      x-microcks-operation:
        delay: 0
        dispatcher: FALLBACK
      security:
      - KeyPair: []
      - ExternalOAuth: []
      - SnowflakeOAuth: []
components:
  schemas:
    StreamingCompleteResponseDataEvent:
      type: object
      description: Streaming text-completion response event.
      properties:
        choices:
          type: array
          items:
            type: object
            properties:
              delta:
                $ref: '#/components/schemas/StreamingCompleteResponseDelta'
          example: []
    GetModelsResponse:
      type: object
      properties:
        models:
          type: array
          items:
            type: string
          example: []
    StreamingCompleteResponse:
      type: object
      description: Server-sent events for streaming text-completion updates.
      x-events:
        data:
          $ref: '#/components/schemas/StreamingCompleteResponseDataEvent'
    StreamingCompleteResponseDelta:
      type: object
      required:
      - type
      discriminator:
        propertyName: type
        mapping:
          text: common-cortex-tool.yaml#/components/schemas/StreamingTextContent
          tool_use: common-cortex-tool.yaml#/components/schemas/StreamingToolUse
    GetModelsRequest:
      type: object
      properties:
        models:
          type: array
          items:
            type: string
          example: []
    GuardrailsConfig:
      type:
      - object
      - 'null'
      title: GuardrailsConfig
      description: Guardrails configuration
      properties:
        enabled:
          type: boolean
          description: Controls whether guardrails are enabled
          example: true
        response_when_unsafe:
          type: string
          description: The response when the guardrails model marks the completion as unsafe
          example: Response filtered by Cortex Guard
    CompleteRequest:
      type: object
      description: LLM text completion request.
      properties:
        model:
          description: The model name. See documentation for possible values.
          type: string
          example: example_value
        messages:
          type: array
          items:
            type: object
            properties:
              role:
                type: string
                description: "Indicates the role of the message, one of 'system', 'user' or 'assistant'.\n\nRules:\n  - A 'user' message must be the last message in the list.\n  - If a 'system' message is specified, it must be the first message.\n  - If a 'assistant' message is specified, it must be immediately before a 'user' message in the list.\n\nMultiple 'assistant' and 'user' messages can be specified, but they must alternate in sequence.\n"
                default: user
              content:
                type: string
                description: The text completion prompt, e.g. 'What is a Large Language Model?'.
              content_list:
                type: array
                description: Contents of toolUse and toolResults
                items:
                  discriminator:
                    propertyName: type
                    mapping:
                      text: common-cortex-tool.yaml#/components/schemas/TextContent
                      tool_result: common-cortex-tool.yaml#/components/schemas/ToolResults
                      tool_use: common-cortex-tool.yaml#/components/schemas/ToolUse
            required:
            - content
          minItems: 1
          example: []
        temperature:
          description: Temperature controls the amount of randomness used in response generation. A higher temperature corresponds to more randomness.
          type:
          - number
          - 'null'
          minimum: 0.0
          example: 42.5
        top_p:
          description: Threshold probability for nucleus sampling. A higher top-p value increases the diversity of tokens that the model considers, while a lower value results in more predictable output.
          type: number
          default: 1.0
          minimum: 0.0
          maximum: 1.0
          example: 42.5
        max_tokens:
          description: The maximum number of output tokens to produce. The default value is model-dependent.
          type: integer
          default: 4096
          minimum: 0
          example: 10
        max_output_tokens:
          deprecated: true
          description: Deprecated in favor of "max_tokens", which has identical behavior.
          type:
          - integer
          - 'null'
          example: 10
        response_format:
          type:
          - object
          - 'null'
          description: An object describing response format config for structured-output mode.
          properties:
            type:
              type: string
              enum:
              - json
              description: The response format type (e.g., "json").
            schema:
              type: object
              description: The schema defining the structure of the response. If the `type` is "json", the `schema` field should contain a valid JSON schema.
          example: example_value
        guardrails:
          $ref: '#/components/schemas/GuardrailsConfig'
        tools:
          description: List of tools to be used during tool calling
          type: array
          items:
            $ref: common-cortex-tool.yaml#/components/schemas/Tool
          example: []
        tool_choice:
          $ref: common-cortex-tool.yaml#/components/schemas/ToolChoice
        provisioned_throughput_id:
          type:
          - string
          - 'null'
          description: The provisioned throughput ID to be used with the request.
          example: '500123'
        sf-ml-xp-inflight-prompt-action:
          type: string
          description: Reserved
          example: example_value
        sf-ml-xp-inflight-prompt-client-id:
          type: string
          description: Reserved
          example: '500123'
        sf-ml-xp-inflight-prompt-public-key:
          type: string
          description: Reserved
          example: example_value
        stream:
          type:
          - boolean
          - 'null'
          default: true
          description: Reserved
          example: true
      required:
      - model
      - messages
    StreamingCompleteResponseDataEvent_2:
      type: object
      description: Streaming text-completion response event.
      properties:
        choices:
          type: array
          items:
            type: object
            properties:
              delta:
                $ref: '#/components/schemas/StreamingCompleteResponseDelta_2'
    GetModelsResponse_2:
      type: object
      properties:
        models:
          type: array
          items:
            type: string
    Tool:
      type: object
      description: 'Defines a tool that can be used by the agent.

        Tools provide specific capabilities like data analysis, search, or generic functions.'
      required:
      - tool_spec
      properties:
        tool_spec:
          type: object
          description: Specification of the tool's type, configuration, and input requirements
          required:
          - type
          - name
          properties:
            type:
              type: string
              description: 'The type of tool capability. Can be specialized types like

                ''cortex_analyst_text_to_sql'' or ''generic'' for general-purpose tools.'
              example: generic
            name:
              type: string
              description: 'Unique identifier for referencing this tool instance.

                Used to match with configuration in tool_resources.'
              example: get_weather
            description:
              type: string
              description: Description of the tool to be considered for tool use
            input_schema:
              type: object
              description: 'JSON Schema definition of the expected input parameters for this tool.

                Required for generic tools to specify their input requirements.'
              properties:
                type:
                  type: string
                  description: The type of the input schema object
                  example: object
                properties:
                  type: object
                  description: Definitions of each input parameter
                required:
                  type: array
                  description: List of required input parameter names
                  items:
                    type: string
              additionalProperties: true
              example:
                type: object
                properties:
                  location:
                    type: string
                    description: The city and state, e.g. San Francisco, CA
                required:
                - location
        cache_control:
          $ref: '#/components/schemas/CacheControl'
    CompleteRequest_2:
      type: object
      description: LLM text completion request.
      properties:
        model:
          description: The model name. See documentation for possible values.
          type: string
        anthropic:
          type:
          - object
          - 'null'
          description: Anthropic specific configuration.
          properties:
            thinking:
              type:
              - object
              - 'null'
              description: Thinking configuration.
              properties:
                budget_tokens:
                  type: integer
                  description: The budget for thinking tokens.
                  example: 2048
              required:
              - budget_tokens
        openai:
          type:
          - object
          - 'null'
          description: OpenAI specific configuration.
          properties:
            reasoning:
              type:
              - object
              - 'null'
              description: Reasoning configuration.
              properties:
                effort:
                  type: string
                  description: The effort level for reasoning.
                  enum:
                  - low
                  - medium
                  - high
        messages:
          type: array
          items:
            type: object
            properties:
              role:
                type: string
                description: "Indicates the role of the message, one of 'system', 'user' or 'assistant'.\n\nRules:\n  - A 'user' message must be the last message in the list.\n  - If a 'system' message is specified, it must be the first message.\n  - If a 'assistant' message is specified, it must be immediately before a 'user' message in the list.\n\nMultiple 'assistant' and 'user' messages can be specified, but they must alternate in sequence.\n"
                default: user
              content:
                type: string
                description: The text completion prompt, e.g. 'What is a Large Language Model?'.
              content_list:
                type: array
                description: Contents of text, toolUse, toolResults, or model-specific output
                items:
                  discriminator:
                    propertyName: type
                    mapping:
                      text: '#/components/schemas/TextContent'
                      tool_result: '#/components/schemas/ToolResults'
                      tool_use: '#/components/schemas/ToolUse'
                      image: '#/components/schemas/Image'
                      anthropic: '#/components/schemas/AnthropicOutput'
                      openai: '#/components/schemas/OpenAIOutput'
          minItems: 1
        temperature:
          description: Temperature controls the amount of randomness used in response generation. A higher temperature corresponds to more randomness.
          type:
          - number
          - 'null'
          minimum: 0.0
        top_p:
          description: Threshold probability for nucleus sampling. A higher top-p value increases the diversity of tokens that the model considers, while a lower value results in more predictable output.
          type: number
          default: 1.0
          minimum: 0.0
          maximum: 1.0
        max_tokens:
          description: The maximum number of output tokens to produce. The default value is model-dependent.
          type: integer
          default: 4096
          minimum: 0
        max_output_tokens:
          deprecated: true
          description: Deprecated in favor of "max_tokens", which has identical behavior.
          type:
          - integer
          - 'null'
        response_format:
          type:
          - object
          - 'null'
          description: An object describing response format config for structured-output mode.
          properties:
            type:
              type: string
              enum:
              - json
              description: The response format type (e.g., "json").
            schema:
              type: object
              description: The schema defining the structure of the response. If the `type` is "json", the `schema` field should contain a valid JSON schema.
        guardrails:
          $ref: '#/components/schemas/GuardrailsConfig_2'
        tools:
          description: List of tools to be used during tool calling
          type: array
          items:
            $ref: '#/components/schemas/Tool'
        tool_choice:
          $ref: '#/components/schemas/ToolChoice'
        provisioned_throughput_id:
          type:
          - string
          - 'null'
          description: The provisioned throughput ID to be used with the request.
        sf-ml-xp-inflight-prompt-action:
          type: string
          description: Reserved
        sf-ml-xp-inflight-prompt-client-id:
          type: string
          description: Reserved
        sf-ml-xp-inflight-prompt-public-key:
          type: string
          description: Reserved
        stream:
          type:
          - boolean
          - 'null'
          default: true
          description: Reserved
      required:
      - model
      - messages
    StreamingCompleteResponseDelta_2:
      type: object
      required:
      - type
      discriminator:
        propertyName: type
        mapping:
          text: '#/components/schemas/StreamingTextContent'
          tool_use: '#/components/schemas/StreamingToolUse'
          anthropic: '#/components/schemas/StreamingAnthropicOutput'
          openai: '#/components/schemas/StreamingOpenAIOutput'
    CacheControl:
      description: reserved
      type:
      - object
      - 'null'
      properties:
        type:
          type: string
          description: 'Identifies the type of cache control.

            “ephemeral” is the only supported cache type currently, which has a 5-minute lifetime.

            '
          enum:
          - ephemeral
    GetModelsRequest_2:
      type: object
      properties:
        models:
          type: array
          items:
            type: string
    GuardrailsConfig_2:
      type:
      - object
      - 'null'
      title: GuardrailsConfig
      description: Guardrails configuration
      properties:
        enabled:
          type: boolean
          description: Controls whether guardrails are enabled
        response_when_unsafe:
          type: string
          description: The response when the guardrails model marks the completion as unsafe
          example: Response filtered by Cortex Guard
    ToolChoice:
      type:
      - object
      - 'null'
      description: 'Configures how tools should be selected and used during the interaction.

        Controls whether tool use is automatic, required, or specific tools should be used.'
      required:
      - type
      properties:
        type:
          type: string
          description: 'Determines how tools are selected:

            * auto - Automatic tool selection (default)

            * required - Must use at least one tool

            * tool - Use specific named tools'
          example: required
        name:
          type: array
          description: List of specific tool names to use when type is 'tool'
          example:
          - Analyst1
          - Search1
          items:
            type: string
  securitySchemes:
    KeyPair:
      $ref: common.yaml#/components/securitySchemes/KeyPair
    ExternalOAuth:
      $ref: common.yaml#/components/securitySchemes/ExternalOAuth
    SnowflakeOAuth:
      $ref: common.yaml#/components/securitySchemes/SnowflakeOAuth
    ProgrammaticAccessToken:
      $ref: common.yaml#/components/securitySchemes/ProgrammaticAccessToken
x-refined-from:
- cortex-inference.yaml
- cortex-inference.yaml