Kimi (Moonshot AI) Chat API
The Chat API from Kimi (Moonshot AI) — 1 operation(s) for chat.
The Chat API from Kimi (Moonshot AI) — 1 operation(s) for chat.
openapi: 3.1.0
info:
title: Moonshot AI Chat API
version: 1.0.0
description: Moonshot AI / Kimi 大语言模型服务 API
servers:
- url: https://api.moonshot.cn
description: 生产环境
tags:
- name: Chat
paths:
/v1/chat/completions:
post:
summary: 创建聊天补全
description: 为聊天消息创建补全结果。支持标准聊天、Partial Mode 和 Tool Use(函数调用)。
tags:
- Chat
security:
- bearerAuth: []
requestBody:
required: true
content:
application/json:
schema:
oneOf:
- $ref: '#/components/schemas/KimiK26ChatRequest'
- $ref: '#/components/schemas/KimiK25ChatRequest'
- $ref: '#/components/schemas/KimiK2ChatRequest'
- $ref: '#/components/schemas/KimiK2ThinkingChatRequest'
- $ref: '#/components/schemas/MoonshotV1ChatRequest'
discriminator:
propertyName: model
mapping:
kimi-k2.6: '#/components/schemas/KimiK26ChatRequest'
kimi-k2.5: '#/components/schemas/KimiK25ChatRequest'
kimi-k2-0905-preview: '#/components/schemas/KimiK2ChatRequest'
kimi-k2-0711-preview: '#/components/schemas/KimiK2ChatRequest'
kimi-k2-turbo-preview: '#/components/schemas/KimiK2ChatRequest'
kimi-k2-thinking: '#/components/schemas/KimiK2ThinkingChatRequest'
kimi-k2-thinking-turbo: '#/components/schemas/KimiK2ThinkingChatRequest'
moonshot-v1-8k: '#/components/schemas/MoonshotV1ChatRequest'
moonshot-v1-32k: '#/components/schemas/MoonshotV1ChatRequest'
moonshot-v1-128k: '#/components/schemas/MoonshotV1ChatRequest'
moonshot-v1-auto: '#/components/schemas/MoonshotV1ChatRequest'
moonshot-v1-8k-vision-preview: '#/components/schemas/MoonshotV1ChatRequest'
moonshot-v1-32k-vision-preview: '#/components/schemas/MoonshotV1ChatRequest'
moonshot-v1-128k-vision-preview: '#/components/schemas/MoonshotV1ChatRequest'
responses:
'200':
description: 聊天补全响应
content:
application/json:
schema:
$ref: '#/components/schemas/ChatCompletionResponse'
'400':
description: 请求错误 - 参数无效
content:
application/json:
schema:
$ref: '#/components/schemas/ErrorResponse'
'401':
description: 未授权 - API 密钥无效或缺失
content:
application/json:
schema:
$ref: '#/components/schemas/ErrorResponse'
'500':
description: 服务器错误
content:
application/json:
schema:
$ref: '#/components/schemas/ErrorResponse'
components:
schemas:
KimiK25ChatRequest:
title: kimi-k2.5
allOf:
- $ref: '#/components/schemas/ChatRequestBase'
- type: object
properties:
model:
type: string
description: 模型 ID
enum:
- kimi-k2.5
default: kimi-k2.5
thinking:
type: object
description: '控制模型是否启用思考能力。可选参数,默认值为 {"type": "enabled"}。'
properties:
type:
type: string
enum:
- enabled
- disabled
description: 启用或禁用思考能力
required:
- type
additionalProperties: false
required:
- model
KimiK2ChatRequest:
title: kimi-k2
allOf:
- $ref: '#/components/schemas/ChatRequestBase'
- type: object
properties:
model:
type: string
description: 模型 ID
enum:
- kimi-k2-0905-preview
- kimi-k2-0711-preview
- kimi-k2-turbo-preview
default: kimi-k2-0905-preview
temperature:
type: number
format: float
description: 采样温度,范围 0 到 1。较高的值(如 0.7)使输出更随机,较低的值(如 0.2)使输出更集中和确定。默认值为 0.6。
default: 0.6
minimum: 0
maximum: 1
top_p:
type: number
format: float
description: 另一种采样方法,模型考虑累积概率质量为 top_p 的 Token 结果。例如 0.1 表示仅考虑概率质量前 10% 的 Token。通常建议只修改此参数或 temperature 其中之一。默认值为 1.0。
default: 1.0
minimum: 0
maximum: 1
n:
type: integer
description: 每条输入消息生成的结果数量。默认为 1,不超过 5。当温度非常接近 0 时,只能返回 1 个结果。
default: 1
minimum: 1
maximum: 5
presence_penalty:
type: number
format: float
description: 存在惩罚,范围 -2.0 到 2.0。正值会根据 Token 是否出现在文本中进行惩罚,增加模型讨论新话题的可能性
default: 0
minimum: -2
maximum: 2
frequency_penalty:
type: number
format: float
description: 频率惩罚,范围 -2.0 到 2.0。正值会根据 Token 在文本中的现有频率进行惩罚,降低模型逐字重复相同短语的可能性
default: 0
minimum: -2
maximum: 2
required:
- model
ChatCompletionResponse:
type: object
properties:
id:
type: string
description: 补全结果的唯一标识符
object:
type: string
description: 对象类型
example: chat.completion
created:
type: integer
description: 补全创建时的 Unix 时间戳
model:
type: string
description: 用于补全的模型
choices:
type: array
description: 补全选项列表
items:
type: object
properties:
index:
type: integer
message:
type: object
properties:
role:
type: string
enum:
- assistant
content:
type:
- string
- 'null'
description: 助手的消息内容
tool_calls:
type: array
description: 模型发起的工具调用
items:
type: object
properties:
id:
type: string
type:
type: string
enum:
- function
function:
type: object
properties:
name:
type: string
arguments:
type: string
description: 函数参数的 JSON 字符串
finish_reason:
type: string
enum:
- stop
- length
- tool_calls
usage:
type: object
properties:
prompt_tokens:
type: integer
description: 提示中的 Token 数量
completion_tokens:
type: integer
description: 补全中的 Token 数量
total_tokens:
type: integer
description: 使用的总 Token 数量
ErrorResponse:
type: object
properties:
error:
type: object
properties:
message:
type: string
description: 描述错误原因的错误消息
type:
type: string
description: 错误类型
code:
type: string
description: 错误码
required:
- message
required:
- error
ChatRequestBase:
type: object
properties:
messages:
type: array
description: '包含迄今为止对话的消息列表。每个元素格式为 {"role": "user", "content": "你好"}。role 支持 system、user、assistant 其一,content 不得为空。content 字段可以是 string,也可以是 array[object](用于多模态输入)'
items:
$ref: '#/components/schemas/Message'
max_tokens:
type: integer
deprecated: true
description: 已弃用,请使用 max_completion_tokens
max_completion_tokens:
type: integer
description: 聊天补全生成的最大 Token 数量。如果不给的话,默认给一个不错的整数比如 1024。如果结果达到最大 Token 数而未结束,finish reason 将为 "length";否则为 "stop"。此值为期望返回的 Token 长度,而非输入加输出的总长度。如果输入加 max_completion_tokens 超出模型上下文窗口,将返回 invalid_request_error。
response_format:
type: object
description: '控制模型输出格式。默认值为 {"type": "text"},即纯文本输出。设置为 {"type": "json_object"} 可启用 JSON 模式,确保输出为合法 JSON 对象(需在 prompt 中引导模型输出 JSON 并指定格式)。设置为 {"type": "json_schema"} 可启用 Structured Output,按指定的 JSON Schema 约束输出结构(推荐,需配合 json_schema 字段使用)。如果您在使用 JSON Schema 时遇到校验问题,欢迎到 walle GitHub Issues (https://github.com/MoonshotAI/walle/issues) 提交反馈。'
properties:
type:
type: string
enum:
- text
- json_object
- json_schema
description: 输出格式类型。text:默认,纯文本输出;json_object:保证输出为合法 JSON 对象;json_schema:按指定 JSON Schema 约束输出(推荐,需配合 json_schema 字段使用)
json_schema:
type: object
description: 当 type 为 json_schema 时使用,定义输出应遵循的 JSON Schema
properties:
name:
type: string
description: Schema 名称,用于标识
strict:
type: boolean
default: true
description: 是否严格按 schema 约束输出。默认为 true。为 true 时 schema 需符合 MFJS 规范,不符合会返回错误或 warning;为 false 时仅保证输出为合法 JSON 对象,不强制约束内部结构。
schema:
type: object
description: JSON Schema 对象,定义输出应遵循的结构。需符合 MFJS(Moonshot Flavored JSON Schema)规范。可使用 walle CLI 工具自检:go install github.com/moonshotai/walle/cmd/walle@latest && walle -schema '你的schema' -level strict
additionalProperties: true
required:
- name
- schema
stop:
oneOf:
- type: string
- type: array
items:
type: string
maxItems: 5
default: null
description: 停用词,完全匹配时将停止输出。匹配到的词本身不会被输出。最多允许 5 个字符串,每个不超过 32 字节
stream:
type: boolean
default: false
description: 是否以流式方式返回响应,默认 false
stream_options:
type: object
description: 流式响应选项
properties:
include_usage:
type: boolean
default: false
description: '如果设置,将在 data: [DONE] 消息之前额外发送一个 chunk。该 chunk 的 usage 字段显示整个请求的 Token 使用统计,choices 字段为空数组。其他所有 chunk 也会包含 usage 字段,但值为 null。注意:如果流中断,可能无法收到包含总 Token 用量的最终 chunk'
tools:
type: array
description: 模型可调用的工具列表
items:
$ref: '#/components/schemas/ToolDefinition'
maxItems: 128
prompt_cache_key:
type: string
default: null
description: 用于缓存相似请求的响应以优化缓存命中率。对于 Coding Agent,通常是代表单个会话的 session id 或 task id;退出并恢复会话时应保持不变。对于 Kimi Code Plan,此字段为必填以提高缓存命中率。对于其他多轮对话 Agent,也建议使用此字段
safety_identifier:
type: string
description: 用于检测可能违反使用政策的用户的稳定标识符。应为唯一标识每个用户的字符串。建议对用户名或邮箱进行哈希处理以避免发送可识别信息
required:
- messages
KimiK2ThinkingChatRequest:
title: kimi-k2-thinking
allOf:
- $ref: '#/components/schemas/ChatRequestBase'
- type: object
properties:
model:
type: string
description: 模型 ID
enum:
- kimi-k2-thinking
- kimi-k2-thinking-turbo
default: kimi-k2-thinking
temperature:
type: number
format: float
description: 采样温度,范围 0 到 1。较高的值(如 0.7)使输出更随机,较低的值(如 0.2)使输出更集中和确定。默认值为 1.0。
default: 1.0
minimum: 0
maximum: 1
top_p:
type: number
format: float
description: 另一种采样方法,模型考虑累积概率质量为 top_p 的 Token 结果。例如 0.1 表示仅考虑概率质量前 10% 的 Token。通常建议只修改此参数或 temperature 其中之一。默认值为 1.0。
default: 1.0
minimum: 0
maximum: 1
n:
type: integer
description: 每条输入消息生成的结果数量。默认为 1,不超过 5。当温度非常接近 0 时,只能返回 1 个结果。
default: 1
minimum: 1
maximum: 5
presence_penalty:
type: number
format: float
description: 存在惩罚,范围 -2.0 到 2.0。正值会根据 Token 是否出现在文本中进行惩罚,增加模型讨论新话题的可能性
default: 0
minimum: -2
maximum: 2
frequency_penalty:
type: number
format: float
description: 频率惩罚,范围 -2.0 到 2.0。正值会根据 Token 在文本中的现有频率进行惩罚,降低模型逐字重复相同短语的可能性
default: 0
minimum: -2
maximum: 2
required:
- model
ToolDefinition:
type: object
properties:
type:
type: string
enum:
- function
function:
type: object
properties:
name:
type: string
description: 函数名称。必须符合正则表达式:^[a-zA-Z_][a-zA-Z0-9-_]{2,63}$
pattern: ^[a-zA-Z_][a-zA-Z0-9-_]{2,63}$
description:
type: string
description: 函数功能描述
parameters:
type: object
description: 函数参数,JSON Schema 格式。需符合 [MFJS(Moonshot Flavored JSON Schema)规范](https://github.com/MoonshotAI/walle/blob/main/docs/mfjs-spec.zh.md)。
additionalProperties: true
strict:
type: boolean
default: true
description: 是否严格按 parameters schema 约束工具调用参数的输出。默认为 true。设为 false 时仅保证输出为合法 JSON 对象,不强制约束内部结构。
required:
- name
- parameters
required:
- type
- function
KimiK26ChatRequest:
title: kimi-k2.6
allOf:
- $ref: '#/components/schemas/ChatRequestBase'
- type: object
properties:
model:
type: string
description: 模型 ID
enum:
- kimi-k2.6
default: kimi-k2.6
thinking:
type: object
description: '控制 kimi-k2.6 模型是否启用思考能力,以及是否完整保留多轮对话中的 reasoning_content。可选参数,默认值为 {"type": "enabled"}。'
properties:
type:
type: string
enum:
- enabled
- disabled
description: 启用或禁用思考能力
keep:
type:
- string
- 'null'
enum:
- all
- null
description: '控制是否保留历史对话轮次(previous turns)的 reasoning_content,从而启用 Preserved Thinking。默认为 `null`,即不保留历史轮次的思考内容。
- `null`(默认)或不传:服务端会忽略历史 turns 的 reasoning_content。
- `"all"`:保留历史 turns 的 reasoning_content 并随上下文一同提供给模型,启用 Preserved Thinking。使用时需把每一轮历史 assistant 消息中的 reasoning_content 原样保留在 messages 中。推荐与 `type: "enabled"` 搭配使用。
- 注意:该参数只影响历史轮次的 reasoning_content;不改变模型在当前 turn 内是否产生/输出思考内容(由 `type` 控制)。关于使用方式的最佳实践,详见 [Preserved Thinking](/guide/use-kimi-k2-thinking-model#preserved-thinking)。'
required:
- type
additionalProperties: false
required:
- model
MoonshotV1ChatRequest:
title: moonshot-v1
allOf:
- $ref: '#/components/schemas/ChatRequestBase'
- type: object
properties:
model:
type: string
description: 模型 ID
enum:
- moonshot-v1-8k
- moonshot-v1-32k
- moonshot-v1-128k
- moonshot-v1-auto
- moonshot-v1-8k-vision-preview
- moonshot-v1-32k-vision-preview
- moonshot-v1-128k-vision-preview
default: moonshot-v1-128k
temperature:
type: number
format: float
description: 采样温度,范围 0 到 1。较高的值(如 0.7)使输出更随机,较低的值(如 0.2)使输出更集中和确定。默认值为 0.0。
default: 0.0
minimum: 0
maximum: 1
top_p:
type: number
format: float
description: 另一种采样方法,模型考虑累积概率质量为 top_p 的 Token 结果。例如 0.1 表示仅考虑概率质量前 10% 的 Token。通常建议只修改此参数或 temperature 其中之一。默认值为 1.0。
default: 1.0
minimum: 0
maximum: 1
n:
type: integer
description: 每条输入消息生成的结果数量。默认为 1,不超过 5。当温度非常接近 0 时,只能返回 1 个结果。
default: 1
minimum: 1
maximum: 5
presence_penalty:
type: number
format: float
description: 存在惩罚,范围 -2.0 到 2.0。正值会根据 Token 是否出现在文本中进行惩罚,增加模型讨论新话题的可能性
default: 0
minimum: -2
maximum: 2
frequency_penalty:
type: number
format: float
description: 频率惩罚,范围 -2.0 到 2.0。正值会根据 Token 在文本中的现有频率进行惩罚,降低模型逐字重复相同短语的可能性
default: 0
minimum: -2
maximum: 2
required:
- model
Message:
type: object
properties:
role:
type: string
enum:
- system
- user
- assistant
description: 消息发送者的角色
content:
oneOf:
- type: string
- type: array
items:
oneOf:
- title: text
type: object
properties:
type:
type: string
enum:
- text
text:
type: string
required:
- type
- text
- title: image_url
type: object
properties:
type:
type: string
enum:
- image_url
image_url:
oneOf:
- type: object
properties:
url:
type: string
required:
- url
- type: string
required:
- type
- image_url
- title: video_url
type: object
properties:
type:
type: string
enum:
- video_url
video_url:
oneOf:
- type: object
properties:
url:
type: string
required:
- url
- type: string
required:
- type
- video_url
description: 消息内容。可以是纯文本字符串,也可以是包含 text/image_url/video_url 类型的对象数组(用于多模态输入)
name:
type: string
default: null
description: 消息发送者的名称(可选)
partial:
type: boolean
default: false
description: 在最后一条 assistant 消息中设置为 true 以启用 Partial Mode
required:
- role
- content
securitySchemes:
bearerAuth:
type: http
scheme: bearer
description: Authorization 请求头需要一个 Bearer 令牌。使用 MOONSHOT_API_KEY 作为令牌。这是一个服务端密钥,请在 [API 密钥页面](https://platform.kimi.com/console/api-keys) 生成。