NVIDIA NIM · OpenAPI Overlay 1.0.0

API Evangelist conversational phrasing for NVIDIA NIM Completions Chat API

2 actions 2 updates phrasing extends openapi/nvidia-nim-chat-api-openapi.yml
Generated by API Evangelist Written by API Evangelist tooling for NVIDIA NIM's API. It is a proposal applied on top of the contract, not a document NVIDIA NIM publishes.
View Overlay File View on GitHub Overlay Specification

What the actions change

x-apievangelist-phrasing

Targets 2

$.info
$.paths['/v1/chat/completions'].post

OpenAPI Overlay

Raw ↑
# Generated by API Evangelist (build-phrasing.py). Our phrasing, not observed demand.
overlay: 1.0.0
info:
  title: API Evangelist conversational phrasing for NVIDIA NIM Completions Chat API
  version: 1.0.0
extends: openapi/nvidia-nim-chat-api-openapi.yml
actions:
- target: $.info
  update:
    x-apievangelist-phrasing:
      method: generated
      generated: '2026-10-01'
      generator: build-phrasing.py
      label: Generated by API Evangelist
      operations: 1
- target: $.paths['/v1/chat/completions'].post
  update:
    x-apievangelist-phrasing:
      intent: Generate a chat reply from a conversation
      effect: write
      questions:
      - Can I send a list of chat messages to a Llama or Nemotron model and get the assistant's reply back?
      - Does the chat completions endpoint support streaming responses and tool or function calling?
      - Can I force a chat model to answer in JSON mode with structured output?
      - How do I send an image inside the messages to a vision-language chat model?
      instructions:
      - text: Send the conversation {messages} to the {model} chat model and return its reply.
        slots:
          messages: requestBody.messages
          model: requestBody.model
      - text: Stream a chat response from {model} for {messages}.
        slots:
          model: requestBody.model
          messages: requestBody.messages
      - text: Ask {model} to answer {messages}, letting it call the tools {tools}.
        slots:
          model: requestBody.model
          messages: requestBody.messages
          tools: requestBody.tools
      - text: Get a chat reply from {model} for {messages} capped at {max_tokens} tokens with temperature {temperature}.
        slots:
          model: requestBody.model
          messages: requestBody.messages
          max_tokens: requestBody.max_tokens
          temperature: requestBody.temperature
      method: generated
      generated: '2026-10-01'