> ## Documentation Index
> Fetch the complete documentation index at: https://dify-6c0370d8-fix-template-upload-size-guidance.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Convert Text to Audio

> **Available for**: Chatflow, Workflow, Agent, Chatbot, Legacy Agent, Text Generator apps.

Converts text to speech audio. Pass `text` to synthesize arbitrary text, or `message_id` to voice an existing message's answer.



## OpenAPI

````yaml /en/api-reference/openapi_service.json post /text-to-audio
openapi: 3.0.1
info:
  description: >-
    REST API for Dify applications and knowledge bases. Application endpoints
    authenticate with an app API key; knowledge endpoints authenticate with a
    dataset API key.
  title: Dify Service API
  version: 1.0.0
servers:
  - description: >-
      Base URL of the Dify Service API. For self-hosted deployments, replace it
      with your own API base URL.
    url: https://{api_base_url}
    variables:
      api_base_url:
        default: api.dify.ai/v1
        description: Host and path of the API base URL, without the `https://` prefix.
security:
  - ApiKeyAuth: []
tags:
  - description: Operations related to chat messages and interactions.
    name: Chat Messages
  - description: File upload and preview operations.
    name: Files
  - description: Operations related to end user information.
    name: End Users
  - description: User feedback operations.
    name: Feedback
  - description: Operations related to managing conversations.
    name: Conversations
  - description: Text-to-Speech and Speech-to-Text operations.
    name: Audio
  - description: Operations to retrieve application settings and information.
    name: Applications
  - description: Operations related to managing annotations for direct replies.
    name: Annotations
  - description: Endpoints for resuming paused workflows that require human input.
    name: Human Input
  - description: Operations for executing and managing workflows.
    name: Workflow Runs
  - description: Operations related to text generation and completion.
    name: Completion Messages
  - description: >-
      Operations for managing knowledge bases, including creation,
      configuration, and retrieval.
    name: Knowledge Bases
  - description: >-
      Operations for creating, updating, and managing documents within a
      knowledge base.
    name: Documents
  - description: Operations for managing document chunks and child chunks.
    name: Chunks
  - description: >-
      Operations for managing knowledge base metadata fields and document
      metadata values.
    name: Metadata
  - description: Operations for managing knowledge base tags and tag bindings.
    name: Tags
  - description: Operations for retrieving available models.
    name: Models
  - description: >-
      Operations for managing and running knowledge pipelines, including
      datasource plugins and pipeline execution.
    name: Knowledge Pipeline
paths:
  /text-to-audio:
    post:
      tags:
        - Audio
      summary: Convert Text to Audio
      description: >-
        **Available for**: Chatflow, Workflow, Agent, Chatbot, Legacy Agent,
        Text Generator apps.


        Converts text to speech audio. Pass `text` to synthesize arbitrary text,
        or `message_id` to voice an existing message's answer.
      operationId: textToAudioChat
      requestBody:
        content:
          application/json:
            examples:
              textToAudioExample:
                summary: Request Example
                value:
                  streaming: false
                  text: Hello, welcome to our service.
                  user: abc-123
                  voice: alloy
            schema:
              $ref: '#/components/schemas/TextToAudioRequest'
        required: true
      responses:
        '200':
          content:
            audio/aac:
              schema:
                format: binary
                type: string
            audio/flac:
              schema:
                format: binary
                type: string
            audio/mp4:
              schema:
                format: binary
                type: string
            audio/mpeg:
              schema:
                format: binary
                type: string
            audio/ogg:
              schema:
                format: binary
                type: string
            audio/wav:
              schema:
                format: binary
                type: string
            audio/webm:
              schema:
                format: binary
                type: string
          description: >-
            Returns the generated audio. The `Content-Type` header reflects the
            provider's audio container, verified from the response bytes when
            recognizable.


            The body can be AAC, FLAC, MP4, MP3, Ogg, WAV, or WebM. Output that
            cannot be recognized is labeled with the provider's declared type,
            or `audio/mpeg` when none is declared.


            Streamed provider output is delivered with chunked transfer
            encoding; the request `streaming` field does not control this.
        '400':
          content:
            application/json:
              examples:
                app_unavailable:
                  summary: app_unavailable
                  value:
                    code: app_unavailable
                    message: App unavailable, please check your app configurations.
                    status: 400
                completion_request_error:
                  summary: completion_request_error
                  value:
                    code: completion_request_error
                    message: Completion request failed.
                    status: 400
                invalid_param:
                  summary: invalid_param
                  value:
                    code: invalid_param
                    message: TTS is not enabled
                    status: 400
                model_currently_not_support:
                  summary: model_currently_not_support
                  value:
                    code: model_currently_not_support
                    message: >-
                      Dify Hosted OpenAI trial currently not support the GPT-4
                      model.
                    status: 400
                provider_not_initialize:
                  summary: provider_not_initialize
                  value:
                    code: provider_not_initialize
                    message: >-
                      No valid model provider credentials found. Please go to
                      Settings -> Model Provider to complete your provider
                      credentials.
                    status: 400
                provider_quota_exceeded:
                  summary: provider_quota_exceeded
                  value:
                    code: provider_quota_exceeded
                    message: >-
                      Your quota for Dify Hosted OpenAI has been exhausted.
                      Please go to Settings -> Model Provider to complete your
                      own provider credentials.
                    status: 400
          description: >-
            - `app_unavailable` : The app is unavailable or misconfigured.

            - `invalid_param` : Text-to-speech is not enabled, `text` is
            missing, or no voice is available.

            - `provider_not_initialize` : No valid model provider credentials
            are configured.

            - `provider_quota_exceeded` : The model provider quota is exhausted.

            - `model_currently_not_support` : The current model does not support
            this operation.

            - `completion_request_error` : The text-to-speech request failed.
        '500':
          content:
            application/json:
              examples:
                internal_server_error:
                  summary: internal_server_error
                  value:
                    code: internal_server_error
                    message: >-
                      The server encountered an internal error and was unable to
                      complete your request. Either the server is overloaded or
                      there is an error in the application.
                    status: 500
          description: '`internal_server_error` : Internal server error.'
components:
  schemas:
    TextToAudioRequest:
      description: >-
        Request body for text-to-audio conversion. Provide either `message_id`
        or `text`.
      properties:
        message_id:
          description: >-
            ID of the message whose answer to voice. Takes priority over `text`
            when both are provided. Get message IDs from [List Conversation
            Messages](/en/api-reference/conversations/list-conversation-messages).
          format: uuid
          type: string
        streaming:
          description: >-
            Accepted for backward compatibility but has no effect. Whether the
            audio is streamed is determined by the configured TTS provider's
            output, not by this field.
          type: boolean
        text:
          description: Text to synthesize into speech.
          type: string
        user:
          description: >-
            End-user identifier, defined by your app and unique within it. See
            [End User Identity](/en/api-reference/guides/end-user-identity).
          type: string
        voice:
          description: >-
            Voice to use for text-to-speech. Available voices depend on the TTS
            provider configured for this app. Use the `voice` value from [Get
            App Parameters](/en/api-reference/applications/get-app-parameters) →
            `text_to_speech.voice` for the default.
          type: string
      type: object
  securitySchemes:
    ApiKeyAuth:
      bearerFormat: API_KEY
      description: >-
        Every request authenticates with an API key: `Authorization: Bearer
        {API_KEY}`. App endpoints take an app API key; knowledge endpoints take
        a knowledge base API key ([Get
        Started](/en/api-reference/guides/get-started)).


        Keep keys server-side; never embed them in client code. Requests with a
        missing or invalid key fail with HTTP `401` (`unauthorized`).
      scheme: bearer
      type: http

````