Dify Audio API

Text-to-Speech and Speech-to-Text operations. 2 operation(s) from the Dify Service API.

Operations 2

POST /audio-to-text Convert Audio to Text #
POST /text-to-audio Convert Text to Audio #

Work with this as data

Every API here is available over the APIs.io API and to AI agents over MCP.

MCP server

One button, every client — Claude, Cursor, VS Code and the rest.

https://apis.io/mcp

Tools for apis

7 MCP tools reach this
  • find_apisBrowse and filter every API in the catalog.
  • get_api_artifactsOne API's artifacts, grouped by type.
  • get_openapiThe primary OpenAPI for this API.
  • find_similar_apisAPIs that look like this one.
  • apis_io_searchSTART HERE — APIs, providers and tags for one query, each with its total.
  • resolveTurn a domain, URL or GitHub org into the provider it belongs to.
  • find_cohortsEvery scored population of providers in the catalog.
All 92 tools →

Call it yourself

curl for this page
This API
curl "https://apis.io/api/v1/apis/dify-audio-api"
All apis
curl "https://apis.io/api/v1/apis?limit=25"

Discovery needs no key. Ratings and market analysis are Pro.

Get an API key

Free tier, no form to fill in. Signing in shares your email address with us — we store it to create your key and to recognise you if you sign in with another provider. See our Privacy Policy and Terms.

A second provider on the same verified email joins the account you already have.

OpenAPI Specification

dify-audio-api-openapi.yml Raw ↑
openapi: 3.0.1
info:
  title: Dify Audio API
  description: REST API for Dify applications and knowledge bases. Application endpoints authenticate
    with an app API key; knowledge endpoints authenticate with a dataset API key.
  version: 1.0.0
servers:
- url: https://{api_base_url}
  description: Base URL of the Dify Service API. For self-hosted deployments, replace it with your own
    API base URL.
  variables:
    api_base_url:
      default: api.dify.ai/v1
      description: Host and path of the API base URL, without the `https://` prefix.
security:
- ApiKeyAuth: []
tags:
- name: Audio
  description: Text-to-Speech and Speech-to-Text operations.
paths:
  /audio-to-text:
    post:
      summary: Convert Audio to Text
      description: '**Available for**: Chatflow, Workflow, Agent, Chatbot, Legacy Agent, Text Generator
        apps.


        Transcribes an uploaded audio file to text using the workspace''s default speech-to-text model.'
      operationId: audioToText
      tags:
      - Audio
      requestBody:
        required: true
        content:
          multipart/form-data:
            schema:
              $ref: '#/components/schemas/AudioToTextRequest'
      responses:
        '200':
          description: Successfully converted audio to text.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/AudioToTextResponse'
              examples:
                audioToTextSuccess:
                  summary: Response Example
                  value:
                    text: Hello, I would like to know more about the iPhone 13 Pro Max.
        '400':
          description: '- `no_audio_uploaded` : No audio file was provided in the `file` field.

            - `speech_to_text_disabled` : Speech-to-text is disabled for this app.

            - `provider_not_support_speech_to_text` : The model provider does not support speech-to-text.

            - `provider_not_initialize` : No valid model provider credentials are configured.

            - `completion_request_error` : The speech recognition request failed.'
          content:
            application/json:
              examples:
                no_audio_uploaded:
                  summary: no_audio_uploaded
                  value:
                    status: 400
                    code: no_audio_uploaded
                    message: Please upload your audio.
                speech_to_text_disabled:
                  summary: speech_to_text_disabled
                  value:
                    status: 400
                    code: speech_to_text_disabled
                    message: Speech to text is disabled.
                provider_not_support_speech_to_text:
                  summary: provider_not_support_speech_to_text
                  value:
                    status: 400
                    code: provider_not_support_speech_to_text
                    message: Provider not support speech to text.
                provider_not_initialize:
                  summary: provider_not_initialize
                  value:
                    status: 400
                    code: provider_not_initialize
                    message: No valid model provider credentials found. Please go to Settings -> Model
                      Provider to complete your provider credentials.
                completion_request_error:
                  summary: completion_request_error
                  value:
                    status: 400
                    code: completion_request_error
                    message: Completion request failed.
        '413':
          description: '`audio_too_large` : The audio file exceeds the `30 MB` size limit.'
          content:
            application/json:
              examples:
                audio_too_large:
                  summary: audio_too_large
                  value:
                    status: 413
                    code: audio_too_large
                    message: Audio size larger than 30 mb
        '415':
          description: '`unsupported_audio_type` : The file''s MIME type is not one of the accepted audio
            types (see the `file` field).'
          content:
            application/json:
              examples:
                unsupported_audio_type:
                  summary: unsupported_audio_type
                  value:
                    status: 415
                    code: unsupported_audio_type
                    message: Audio type not allowed.
        '500':
          description: '`internal_server_error` : Internal server error.'
          content:
            application/json:
              examples:
                internal_server_error:
                  summary: internal_server_error
                  value:
                    status: 500
                    code: internal_server_error
                    message: The server encountered an internal error and was unable to complete your
                      request. Either the server is overloaded or there is an error in the application.
      x-mint:
        href: /en/api-reference/audio/convert-audio-to-text
        metadata:
          title: Convert Audio to Text
          sidebarTitle: Convert Audio to Text
  /text-to-audio:
    post:
      summary: Convert Text to Audio
      description: '**Available for**: Chatflow, Workflow, Agent, Chatbot, Legacy Agent, Text Generator
        apps.


        Converts text to speech audio. Pass `text` to synthesize arbitrary text, or `message_id` to voice
        an existing message''s answer.'
      operationId: textToAudioChat
      tags:
      - Audio
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/TextToAudioRequest'
            examples:
              textToAudioExample:
                summary: Request Example
                value:
                  text: Hello, welcome to our service.
                  user: abc-123
                  voice: alloy
                  streaming: false
      responses:
        '200':
          description: 'Returns the generated audio. The `Content-Type` header reflects the provider''s
            audio container, verified from the response bytes when recognizable.


            The body can be AAC, FLAC, MP4, MP3, Ogg, WAV, or WebM. Output that cannot be recognized is
            labeled with the provider''s declared type, or `audio/mpeg` when none is declared.


            Streamed provider output is delivered with chunked transfer encoding; the request `streaming`
            field does not control this.'
          content:
            audio/aac:
              schema:
                type: string
                format: binary
            audio/flac:
              schema:
                type: string
                format: binary
            audio/mp4:
              schema:
                type: string
                format: binary
            audio/mpeg:
              schema:
                type: string
                format: binary
            audio/ogg:
              schema:
                type: string
                format: binary
            audio/wav:
              schema:
                type: string
                format: binary
            audio/webm:
              schema:
                type: string
                format: binary
        '400':
          description: '- `app_unavailable` : The app is unavailable or misconfigured.

            - `invalid_param` : Text-to-speech is not enabled, `text` is missing, or no voice is available.

            - `provider_not_initialize` : No valid model provider credentials are configured.

            - `provider_quota_exceeded` : The model provider quota is exhausted.

            - `model_currently_not_support` : The current model does not support this operation.

            - `completion_request_error` : The text-to-speech request failed.'
          content:
            application/json:
              examples:
                app_unavailable:
                  summary: app_unavailable
                  value:
                    status: 400
                    code: app_unavailable
                    message: App unavailable, please check your app configurations.
                invalid_param:
                  summary: invalid_param
                  value:
                    status: 400
                    code: invalid_param
                    message: TTS is not enabled
                provider_not_initialize:
                  summary: provider_not_initialize
                  value:
                    status: 400
                    code: provider_not_initialize
                    message: No valid model provider credentials found. Please go to Settings -> Model
                      Provider to complete your provider credentials.
                provider_quota_exceeded:
                  summary: provider_quota_exceeded
                  value:
                    status: 400
                    code: provider_quota_exceeded
                    message: Your quota for Dify Hosted OpenAI has been exhausted. Please go to Settings
                      -> Model Provider to complete your own provider credentials.
                model_currently_not_support:
                  summary: model_currently_not_support
                  value:
                    status: 400
                    code: model_currently_not_support
                    message: Dify Hosted OpenAI trial currently not support the GPT-4 model.
                completion_request_error:
                  summary: completion_request_error
                  value:
                    status: 400
                    code: completion_request_error
                    message: Completion request failed.
        '500':
          description: '`internal_server_error` : Internal server error.'
          content:
            application/json:
              examples:
                internal_server_error:
                  summary: internal_server_error
                  value:
                    status: 500
                    code: internal_server_error
                    message: The server encountered an internal error and was unable to complete your
                      request. Either the server is overloaded or there is an error in the application.
      x-mint:
        href: /en/api-reference/audio/convert-text-to-audio
        metadata:
          title: Convert Text to Audio
          sidebarTitle: Convert Text to Audio
components:
  schemas:
    AudioToTextRequest:
      type: object
      description: Request body for audio-to-text conversion.
      required:
      - file
      properties:
        file:
          type: string
          format: binary
          description: 'Audio file to transcribe. Accepted MIME types: `audio/mp3`, `audio/m4a` (also
            accepted as `audio/x-m4a`), `audio/wav`, `audio/amr`, `audio/mpga`. Other types, including
            the common `audio/mpeg`, are rejected with `unsupported_audio_type`. Maximum size `30 MB`.'
        user:
          type: string
          description: End-user identifier, defined by your app and unique within it. See [End User Identity](/en/api-reference/guides/end-user-identity).
    AudioToTextResponse:
      type: object
      properties:
        text:
          type: string
          description: Output text from speech recognition.
    TextToAudioRequest:
      type: object
      description: Request body for text-to-audio conversion. Provide either `message_id` or `text`.
      properties:
        message_id:
          type: string
          format: uuid
          description: ID of the message whose answer to voice. Takes priority over `text` when both are
            provided. Get message IDs from [List Conversation Messages](/en/api-reference/conversations/list-conversation-messages).
        text:
          type: string
          description: Text to synthesize into speech.
        user:
          type: string
          description: End-user identifier, defined by your app and unique within it. See [End User Identity](/en/api-reference/guides/end-user-identity).
        voice:
          type: string
          description: Voice to use for text-to-speech. Available voices depend on the TTS provider configured
            for this app. Use the `voice` value from [Get App Parameters](/en/api-reference/applications/get-app-parameters)
            → `text_to_speech.voice` for the default.
        streaming:
          type: boolean
          description: Accepted for backward compatibility but has no effect. Whether the audio is streamed
            is determined by the configured TTS provider's output, not by this field.
  securitySchemes:
    ApiKeyAuth:
      type: http
      scheme: bearer
      bearerFormat: API_KEY
      description: 'Every request authenticates with an API key: `Authorization: Bearer {API_KEY}`. App
        endpoints take an app API key; knowledge endpoints take a knowledge base API key ([Get Started](/en/api-reference/guides/get-started)).


        Keep keys server-side; never embed them in client code. Requests with a missing or invalid key
        fail with HTTP `401` (`unauthorized`).'
x-provenance:
  generated: '2026-09-06'
  method: derived
  source: openapi/_original/dify-service-api-openapi.json
  note: Per-tag split of the first-party Dify Service API OpenAPI harvested from https://docs.dify.ai/en/api-reference/openapi_service.json
    (advertised in https://docs.dify.ai/llms.txt). Paths, schemas and operationIds are verbatim from that
    spec.