Dify Audio API
Text-to-Speech and Speech-to-Text operations. 2 operation(s) from the Dify Service API.
Text-to-Speech and Speech-to-Text operations. 2 operation(s) from the Dify Service API.
Every API here is available over the APIs.io API and to AI agents over MCP.
One button, every client — Claude, Cursor, VS Code and the rest.
https://apis.io/mcp
find_apisBrowse and filter every API in the catalog.get_api_artifactsOne API's artifacts, grouped by type.get_openapiThe primary OpenAPI for this API.find_similar_apisAPIs that look like this one.apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.resolveTurn a domain, URL or GitHub org into the provider it belongs to.find_cohortsEvery scored population of providers in the catalog.curl "https://apis.io/api/v1/apis/dify-audio-api"
curl "https://apis.io/api/v1/apis?limit=25"
Discovery needs no key. Ratings and market analysis are Pro.
Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.
A second provider on the same verified email joins the account you already have.
openapi: 3.0.1
info:
title: Dify Audio API
description: REST API for Dify applications and knowledge bases. Application endpoints authenticate
with an app API key; knowledge endpoints authenticate with a dataset API key.
version: 1.0.0
servers:
- url: https://{api_base_url}
description: Base URL of the Dify Service API. For self-hosted deployments, replace it with your own
API base URL.
variables:
api_base_url:
default: api.dify.ai/v1
description: Host and path of the API base URL, without the `https://` prefix.
security:
- ApiKeyAuth: []
tags:
- name: Audio
description: Text-to-Speech and Speech-to-Text operations.
paths:
/audio-to-text:
post:
summary: Convert Audio to Text
description: '**Available for**: Chatflow, Workflow, Agent, Chatbot, Legacy Agent, Text Generator
apps.
Transcribes an uploaded audio file to text using the workspace''s default speech-to-text model.'
operationId: audioToText
tags:
- Audio
requestBody:
required: true
content:
multipart/form-data:
schema:
$ref: '#/components/schemas/AudioToTextRequest'
responses:
'200':
description: Successfully converted audio to text.
content:
application/json:
schema:
$ref: '#/components/schemas/AudioToTextResponse'
examples:
audioToTextSuccess:
summary: Response Example
value:
text: Hello, I would like to know more about the iPhone 13 Pro Max.
'400':
description: '- `no_audio_uploaded` : No audio file was provided in the `file` field.
- `speech_to_text_disabled` : Speech-to-text is disabled for this app.
- `provider_not_support_speech_to_text` : The model provider does not support speech-to-text.
- `provider_not_initialize` : No valid model provider credentials are configured.
- `completion_request_error` : The speech recognition request failed.'
content:
application/json:
examples:
no_audio_uploaded:
summary: no_audio_uploaded
value:
status: 400
code: no_audio_uploaded
message: Please upload your audio.
speech_to_text_disabled:
summary: speech_to_text_disabled
value:
status: 400
code: speech_to_text_disabled
message: Speech to text is disabled.
provider_not_support_speech_to_text:
summary: provider_not_support_speech_to_text
value:
status: 400
code: provider_not_support_speech_to_text
message: Provider not support speech to text.
provider_not_initialize:
summary: provider_not_initialize
value:
status: 400
code: provider_not_initialize
message: No valid model provider credentials found. Please go to Settings -> Model
Provider to complete your provider credentials.
completion_request_error:
summary: completion_request_error
value:
status: 400
code: completion_request_error
message: Completion request failed.
'413':
description: '`audio_too_large` : The audio file exceeds the `30 MB` size limit.'
content:
application/json:
examples:
audio_too_large:
summary: audio_too_large
value:
status: 413
code: audio_too_large
message: Audio size larger than 30 mb
'415':
description: '`unsupported_audio_type` : The file''s MIME type is not one of the accepted audio
types (see the `file` field).'
content:
application/json:
examples:
unsupported_audio_type:
summary: unsupported_audio_type
value:
status: 415
code: unsupported_audio_type
message: Audio type not allowed.
'500':
description: '`internal_server_error` : Internal server error.'
content:
application/json:
examples:
internal_server_error:
summary: internal_server_error
value:
status: 500
code: internal_server_error
message: The server encountered an internal error and was unable to complete your
request. Either the server is overloaded or there is an error in the application.
x-mint:
href: /en/api-reference/audio/convert-audio-to-text
metadata:
title: Convert Audio to Text
sidebarTitle: Convert Audio to Text
/text-to-audio:
post:
summary: Convert Text to Audio
description: '**Available for**: Chatflow, Workflow, Agent, Chatbot, Legacy Agent, Text Generator
apps.
Converts text to speech audio. Pass `text` to synthesize arbitrary text, or `message_id` to voice
an existing message''s answer.'
operationId: textToAudioChat
tags:
- Audio
requestBody:
required: true
content:
application/json:
schema:
$ref: '#/components/schemas/TextToAudioRequest'
examples:
textToAudioExample:
summary: Request Example
value:
text: Hello, welcome to our service.
user: abc-123
voice: alloy
streaming: false
responses:
'200':
description: 'Returns the generated audio. The `Content-Type` header reflects the provider''s
audio container, verified from the response bytes when recognizable.
The body can be AAC, FLAC, MP4, MP3, Ogg, WAV, or WebM. Output that cannot be recognized is
labeled with the provider''s declared type, or `audio/mpeg` when none is declared.
Streamed provider output is delivered with chunked transfer encoding; the request `streaming`
field does not control this.'
content:
audio/aac:
schema:
type: string
format: binary
audio/flac:
schema:
type: string
format: binary
audio/mp4:
schema:
type: string
format: binary
audio/mpeg:
schema:
type: string
format: binary
audio/ogg:
schema:
type: string
format: binary
audio/wav:
schema:
type: string
format: binary
audio/webm:
schema:
type: string
format: binary
'400':
description: '- `app_unavailable` : The app is unavailable or misconfigured.
- `invalid_param` : Text-to-speech is not enabled, `text` is missing, or no voice is available.
- `provider_not_initialize` : No valid model provider credentials are configured.
- `provider_quota_exceeded` : The model provider quota is exhausted.
- `model_currently_not_support` : The current model does not support this operation.
- `completion_request_error` : The text-to-speech request failed.'
content:
application/json:
examples:
app_unavailable:
summary: app_unavailable
value:
status: 400
code: app_unavailable
message: App unavailable, please check your app configurations.
invalid_param:
summary: invalid_param
value:
status: 400
code: invalid_param
message: TTS is not enabled
provider_not_initialize:
summary: provider_not_initialize
value:
status: 400
code: provider_not_initialize
message: No valid model provider credentials found. Please go to Settings -> Model
Provider to complete your provider credentials.
provider_quota_exceeded:
summary: provider_quota_exceeded
value:
status: 400
code: provider_quota_exceeded
message: Your quota for Dify Hosted OpenAI has been exhausted. Please go to Settings
-> Model Provider to complete your own provider credentials.
model_currently_not_support:
summary: model_currently_not_support
value:
status: 400
code: model_currently_not_support
message: Dify Hosted OpenAI trial currently not support the GPT-4 model.
completion_request_error:
summary: completion_request_error
value:
status: 400
code: completion_request_error
message: Completion request failed.
'500':
description: '`internal_server_error` : Internal server error.'
content:
application/json:
examples:
internal_server_error:
summary: internal_server_error
value:
status: 500
code: internal_server_error
message: The server encountered an internal error and was unable to complete your
request. Either the server is overloaded or there is an error in the application.
x-mint:
href: /en/api-reference/audio/convert-text-to-audio
metadata:
title: Convert Text to Audio
sidebarTitle: Convert Text to Audio
components:
schemas:
AudioToTextRequest:
type: object
description: Request body for audio-to-text conversion.
required:
- file
properties:
file:
type: string
format: binary
description: 'Audio file to transcribe. Accepted MIME types: `audio/mp3`, `audio/m4a` (also
accepted as `audio/x-m4a`), `audio/wav`, `audio/amr`, `audio/mpga`. Other types, including
the common `audio/mpeg`, are rejected with `unsupported_audio_type`. Maximum size `30 MB`.'
user:
type: string
description: End-user identifier, defined by your app and unique within it. See [End User Identity](/en/api-reference/guides/end-user-identity).
AudioToTextResponse:
type: object
properties:
text:
type: string
description: Output text from speech recognition.
TextToAudioRequest:
type: object
description: Request body for text-to-audio conversion. Provide either `message_id` or `text`.
properties:
message_id:
type: string
format: uuid
description: ID of the message whose answer to voice. Takes priority over `text` when both are
provided. Get message IDs from [List Conversation Messages](/en/api-reference/conversations/list-conversation-messages).
text:
type: string
description: Text to synthesize into speech.
user:
type: string
description: End-user identifier, defined by your app and unique within it. See [End User Identity](/en/api-reference/guides/end-user-identity).
voice:
type: string
description: Voice to use for text-to-speech. Available voices depend on the TTS provider configured
for this app. Use the `voice` value from [Get App Parameters](/en/api-reference/applications/get-app-parameters)
→ `text_to_speech.voice` for the default.
streaming:
type: boolean
description: Accepted for backward compatibility but has no effect. Whether the audio is streamed
is determined by the configured TTS provider's output, not by this field.
securitySchemes:
ApiKeyAuth:
type: http
scheme: bearer
bearerFormat: API_KEY
description: 'Every request authenticates with an API key: `Authorization: Bearer {API_KEY}`. App
endpoints take an app API key; knowledge endpoints take a knowledge base API key ([Get Started](/en/api-reference/guides/get-started)).
Keep keys server-side; never embed them in client code. Requests with a missing or invalid key
fail with HTTP `401` (`unauthorized`).'
x-provenance:
generated: '2026-09-06'
method: derived
source: openapi/_original/dify-service-api-openapi.json
note: Per-tag split of the first-party Dify Service API OpenAPI harvested from https://docs.dify.ai/en/api-reference/openapi_service.json
(advertised in https://docs.dify.ai/llms.txt). Paths, schemas and operationIds are verbatim from that
spec.