使用 Gemini TTS 生成语音

Gemini 文字转语音 (TTS) 模型可将文字转换为单人或多人音频。您可以使用结构化对话元数据 (speech_metadata) 和内嵌语音标记来控制音频的风格、口音、语速和语气。

本页介绍了如何使用 Gemini Enterprise API 的 generateContent 和 streamGenerateContent 方法通过以下模型生成语音:

TTS 模型专为精准朗读文本而量身打造,可对风格和声音进行精细控制,适用于旁白、有声读物和语音代理回答等场景。如需进行交互式实时对话,并使用音频输入和输出,请改用 Live API。

如果您使用的是较早的 Gemini TTS 模型(例如 gemini-3.1-flash-tts-preview),请参阅迁移指南。

支持的模型

模型 一位说话者 多说话人 流式 语音设计 语音复刻
Gemini 3.8 Flash TTS (gemini-3.8-flash-tts)
Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts)
Gemini 3.1 Flash TTS 预览版 (gemini-3.1-flash-tts-preview)
Gemini 2.5 Pro TTS (gemini-2.5-pro-tts)
Gemini 2.5 Flash TTS (gemini-2.5-flash-tts)

本页面上的示例和功能适用于 Gemini 3.8 TTS 模型。对于早期模型,请参阅Google Cloud Text-to-Speech 文档中的 Gemini-TTS。

何时使用哪种模型

Gemini 3.8 TTS 模型具有相同的请求架构和提示格式,因此您可以通过更改模型 ID 在它们之间切换:

  • 如果需要优先考虑声音保真度、细致的表演和富有表现力的控制,请使用 Gemini 3.8 Flash TTS (gemini-3.8-flash-tts)。它适用于工作室级创意工作、复杂的多人对话、频繁的语音片段标记、发音困难的词语、地区或少数民族方言,以及需要稳定声音和室内音调的长篇旁白。
  • 使用 Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) 可快速高效地处理生产工作负载。它是 gemini-3.1-flash-tts-preview 的推荐替代方案,适用于大批量批处理制作、对话语音代理级联、朗读功能、语音复制以及主要语言的日常单人语音。

准备工作

  1. 设置项目并启用 Gemini Enterprise API。
  2. 为您的开发环境配置应用默认凭证。
  3. 如需使用 Python 示例,请安装 Google Gen AI SDK 2.25.0 版或更高版本:

    pip install --upgrade "google-genai>=2.25.0"
    

Gemini 3.8 TTS 模型可在 global 位置使用。向 aiplatform.googleapis.com 端点发送请求。

单说话者 TTS

如需将文本转换为单人音频,请在 parts[].text 中传递逐字转写内容,在 parts[].speech_metadata 中添加可选的轮次级样式,并在 speechConfig.voiceConfig.voice 中设置语音。voice 字段接受预建语音名称、扩展语音库语音 ID,或设计或复制的语音的 ID (voice_...)(或可选的无状态 voicekey_... 键)。

默认情况下,一元请求会返回完整的 WAV 文件(16 位 PCM、24 kHz、单声道),因此您可以将音频字节直接写入 .wav 文件。流式传输请求会返回原始 PCM 块。如需了解详情和其他编码,请参阅音频输出格式。

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

response = client.models.generate_content(
    model="gemini-3.8-flash-tts",
    contents=[{
        "role": "user",
        "parts": [{
            "text": "Have a wonderful day!",
            "speech_metadata": {"style": "cheerful and friendly"},
        }],
    }],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {"voice_config": {"voice": "Kore"}},
    },
)

# The SDK has already decoded the base64 audio, so inline_data.data is a
# complete WAV file by default.
with open("out.wav", "wb") as f:
    f.write(response.candidates[0].content.parts[0].inline_data.data)

REST

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
  -d '{
    "contents": [{
      "role": "user",
      "parts": [{
        "text": "Have a wonderful day!",
        "speechMetadata": {"style": "cheerful and friendly"}
      }]
    }],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "voiceConfig": {"voice": "Kore"}
      }
    }
  }' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > out.wav

单说话者流式 TTS

如需在模型仍在合成音频时接收音频,请使用 streamGenerateContent 方法。流式传输响应会返回无标头的原始 16 位 PCM 分块(24 kHz,单声道),因此您可以在每个分块到达时将其传递给播放器、套接字或其他音频流水线。在以下 Python 示例中,emit_audio 函数代表该目的地。

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

def emit_audio(pcm: bytes) -> None:
    """Sends a chunk of raw 16-bit PCM audio (24 kHz, mono) downstream."""
    # Replace with your audio destination, such as a player, a WebSocket,
    # or a telephony stream.
    print(f"Received {len(pcm)} bytes of audio")

response_stream = client.models.generate_content_stream(
    model="gemini-3.8-flash-lite-tts",
    contents=[{
        "role": "user",
        "parts": [{
            "text": "Have a wonderful day!",
            "speech_metadata": {"style": "cheerful and friendly"},
        }],
    }],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {"voice_config": {"voice": "Kore"}},
    },
)

for chunk in response_stream:
    if not chunk.candidates or not chunk.candidates[0].content:
        continue
    for part in chunk.candidates[0].content.parts or []:
        if part.inline_data and part.inline_data.data:
            emit_audio(part.inline_data.data)

REST

curl -N -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  "https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-lite-tts:streamGenerateContent?alt=sse" \
  -d '{
    "contents": [{
      "role": "user",
      "parts": [{
        "text": "Have a wonderful day!",
        "speechMetadata": {"style": "cheerful and friendly"}
      }]
    }],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "voiceConfig": {"voice": "Kore"}
      }
    }
  }' | sed -n 's/^data: //p' \
    | jq -r '.candidates[0].content.parts[0].inlineData.data // empty' \
    | while read -r chunk; do echo "$chunk" | base64 --decode; done > streamed.pcm

streamed.pcm 文件包含原始 16 位 PCM 音频,采样率 24 kHz,单声道。

多说话人 TTS

对于两位讲话者之间的对话,请在 speechConfig.multiSpeakerVoiceConfig.speakerVoiceConfigs 中配置两位讲话者。然后,将每个对话轮次作为单独的 part 传递,将每个对话轮次的 speech_metadata.speaker 设置为配置的发言人名称之一,并添加可选的对话轮次级 style。

多音箱请求需要两个音箱。

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

response = client.models.generate_content(
    model="gemini-3.8-flash-tts",
    contents=[{
        "role": "user",
        "parts": [
            {
                "text": "How's it going today, Jane?",
                "speech_metadata": {"speaker": "Joe", "style": "cheerful and friendly"},
            },
            {
                "text": "Not too bad, how about you? Ready to test these new voices?",
                "speech_metadata": {"speaker": "Jane", "style": "calm and relaxed"},
            },
        ],
    }],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {
            "multi_speaker_voice_config": {
                "speaker_voice_configs": [
                    {"speaker": "Joe", "voice_config": {"voice": "Puck"}},
                    {"speaker": "Jane", "voice_config": {"voice": "Kore"}},
                ]
            }
        },
    },
)

with open("dialogue.wav", "wb") as f:
    f.write(response.candidates[0].content.parts[0].inline_data.data)

REST

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
  -d '{
    "contents": [{
      "role": "user",
      "parts": [
        {
          "text": "How'\''s it going today, Jane?",
          "speechMetadata": {"speaker": "Joe", "style": "cheerful and friendly"}
        },
        {
          "text": "Not too bad, how about you? Ready to test these new voices?",
          "speechMetadata": {"speaker": "Jane", "style": "calm and relaxed"}
        }
      ]
    }],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "multiSpeakerVoiceConfig": {
          "speakerVoiceConfigs": [
            {"speaker": "Joe", "voiceConfig": {"voice": "Puck"}},
            {"speaker": "Jane", "voiceConfig": {"voice": "Kore"}}
          ]
        }
      }
    }
  }' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > dialogue.wav

多说话人流式 TTS

多音箱请求也支持流式传输。使用与多说话人 TTS 中相同的请求,但使用 streamGenerateContent 方法。

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

def emit_audio(pcm: bytes) -> None:
    """Sends a chunk of raw 16-bit PCM audio (24 kHz, mono) downstream."""
    # Replace with your audio destination, such as a player, a WebSocket,
    # or a telephony stream.
    print(f"Received {len(pcm)} bytes of audio")

response_stream = client.models.generate_content_stream(
    model="gemini-3.8-flash-tts",
    contents=[{
        "role": "user",
        "parts": [
            {
                "text": "How's it going today, Jane?",
                "speech_metadata": {"speaker": "Joe", "style": "cheerful and friendly"},
            },
            {
                "text": "Not too bad, how about you? Ready to test these new voices?",
                "speech_metadata": {"speaker": "Jane", "style": "calm and relaxed"},
            },
        ],
    }],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {
            "multi_speaker_voice_config": {
                "speaker_voice_configs": [
                    {"speaker": "Joe", "voice_config": {"voice": "Puck"}},
                    {"speaker": "Jane", "voice_config": {"voice": "Kore"}},
                ]
            }
        },
    },
)

for chunk in response_stream:
    if not chunk.candidates or not chunk.candidates[0].content:
        continue
    for part in chunk.candidates[0].content.parts or []:
        if part.inline_data and part.inline_data.data:
            emit_audio(part.inline_data.data)

REST

curl -N -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  "https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:streamGenerateContent?alt=sse" \
  -d '{
    "contents": [{
      "role": "user",
      "parts": [
        {
          "text": "How'\''s it going today, Jane?",
          "speechMetadata": {"speaker": "Joe", "style": "cheerful and friendly"}
        },
        {
          "text": "Not too bad, how about you? Ready to test these new voices?",
          "speechMetadata": {"speaker": "Jane", "style": "calm and relaxed"}
        }
      ]
    }],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "speechConfig": {
        "multiSpeakerVoiceConfig": {
          "speakerVoiceConfigs": [
            {"speaker": "Joe", "voiceConfig": {"voice": "Puck"}},
            {"speaker": "Jane", "voiceConfig": {"voice": "Kore"}}
          ]
        }
      }
    }
  }' | sed -n 's/^data: //p' \
    | jq -r '.candidates[0].content.parts[0].inlineData.data // empty' \
    | while read -r chunk; do echo "$chunk" | base64 --decode; done > dialogue_streamed.pcm

dialogue_streamed.pcm 文件包含原始 16 位 PCM 音频,采样率 24 kHz,单声道。

使用元数据和标记控制语音风格

Gemini 3.8 TTS 模型将 text 字段视为逐字转写内容。写入文本中的说明(例如 "Say cheerfully: Hello!" 或 "Speaker 1: Hello!")可能会被朗读出来。如需控制交付,请按范围拆分说明:

  • 持续的轮次级交付 (speech_metadata.style):将适用于整个轮次的情感、交付风格、韵律、语速和音量放在 speech_metadata.style 中。例如 "style": "whispered urgently"、"style": "out of breath" 或 "style": "warm and enthusiastic"。
  • 时间点事件(内嵌标记):将短暂的非语音声音和停顿放在转写内容中的尖括号内。例如,"Wait... <short pause> did you hear that? <sigh>" 或 "Excuse me <cough> as I was saying..."。

如需更多指导,请参阅提示指南。

语音选项

Gemini 3.8 TTS 模型支持四种选择或创建语音的方式:

  1. 预建语音:30 种精选语音。
  2. 扩展语音库:包含 2,000 多种精选的预设语音,涵盖多种语言、地区口音和角色人物。
  3. 语音设计:使用 Voices API create 方法(在 Google Gen AI SDK 中为 client.voices.create(),或 POST https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices)根据自然语言说明创建自定义语音。将 type 设置为 VOICE_TYPE_PROMPTED,并将 store 设置为 true。响应会返回存储的 voice_... ID 和语音样本。
  4. 语音复刻:使用相同的 create 方法,根据参考音频和同意书音频复刻说话人的声音。将 type 设置为 VOICE_TYPE_REPLICATED。使用 store: true 时,响应会返回存储的 voice_... ID。使用 store: false 时,它会返回无状态的 voicekey_... 密钥。

自定义语音限制和 TTL

语音类型 存储模式 管理 失效时间
存储的语音(voice_...,设计或复制) store: true Voices API list、get 和 delete 方法 自上次使用起 1 年后
语音复刻键 (voicekey_...) store: false 由您的应用管理。Google 不会保留密钥。 自创建之日起 7 天后

使用已存储的声音生成语音会重新开始计算一年的保留期限,因此您经常使用的声音不会过期。语音创建请求受每项目每分钟配额的限制。如果超出此限制,create 方法会返回 RESOURCE_EXHAUSTED 错误。如需了解详情,请参阅配额和系统限制。

预建语音

下表列出了每个预建语音名称及其特征:

  • Achernar:柔和
  • Achird:友好
  • Algenib:Gravelly
  • Algieba:平滑
  • Alnilam:坚定
  • Aoede:轻快
  • Autonoe:清亮
  • Callirrhoe:随和
  • Charon:信息丰富
  • Despina:流畅
  • Enceladus:气声
  • Erinome:清晰
  • Fenrir:激昂
  • Gacrux:成熟
  • Iapetus:清晰
  • Kore:坚定
  • Laomedeia:欢快
  • Leda:青春活力
  • Orus:坚固
  • Puck:欢快
  • Pulcherrima:直率
  • Rasalgethi:信息丰富
  • Sadachbia:活泼
  • Sadaltager:博学
  • Schedar:均匀
  • Sulafat:暖色调
  • Umbriel:轻松
  • Vindemiatrix:轻柔
  • Zephyr:明快
  • Zubenelgenubi:随性

扩展语音库

除了 30 种预建语音外,扩展语音库还提供 2,000 多种精选的预设语音,涵盖多种语言、地区口音、角色人物和使用情形(例如有声读物、对话式代理和新闻)。 如需使用库语音,请在 speechConfig.voiceConfig.voice 中传递其 ID,传递方式与传递预构建语音名称相同。

如需浏览该库,请调用 Voices API list 方法 (ListVoices)。该方法会返回存储在项目中的自定义声音,然后返回预建声音和扩展声音库声音。每个库声音都包含其 id 和元数据,例如 language_code、accent、gender、pitch、persona 和 description。如需缩小结果范围,请使用以下过滤条件:

  • type:语音来源:prebuilt 表示预建语音和库语音,prompted 表示设计语音,replicated 表示复制语音。 如果您传递多个值,该方法会返回与其中任何一个值匹配的声音。
  • search:要匹配的文本(不区分大小写),用于与每个语音的显示名称和说明进行匹配。

以下示例在库中搜索旁白语音:

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

response = client.voices.list(type_=["prebuilt"], search="narrator", page_size=50)
for voice in response.voices or []:
    print(voice.id, voice.language_code, voice.gender, voice.description)

REST

curl -G \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices \
  --data-urlencode "type=prebuilt" \
  --data-urlencode "search=narrator" \
  --data-urlencode "pageSize=50"

list 方法每页最多返回 50 个声音。如需获取下一页,请将响应中的 next_page_token 值作为 page_token 参数传递。

音频输出格式

默认音频编码取决于方法:

  • 一元请求 (generateContent) 会返回包含 RIFF 标头的完整 WAV 文件 (audio/wav):16 位 PCM、24 kHz、单声道。
  • 流式请求 (streamGenerateContent) 返回无标头的原始 16 位有符号小端字节序 PCM 块 (audio/l16),采样率为 24 kHz,单声道。

如需请求其他编码,请将 generationConfig.responseFormat 设置为音频格式。Gemini 3.8 TTS 模型支持以下 mimeType 值:

mimeType 值 格式 说明
AUDIO_WAV (一元默认) WAV (audio/wav) 包含 RIFF 标头的完整 WAV 文件(16 位 PCM、24 kHz、单声道)。
AUDIO_L16 (流式默认) 线性 PCM (audio/l16) 无标头的原始 16 位有符号小端序 PCM。用于流式传输、自定义音频流水线或串联多轮对话片段。
AUDIO_MULAW μ-law (audio/mulaw) G.711 μ-law 压扩音频(8 kHz,单声道),常用于北美和日本的电话。
AUDIO_ALAW A-law (audio/alaw) G.711 A-law 压缩音频(8 kHz,单声道),通常用于欧洲和国际电话。

模型会忽略 sampleRate 字段。线性 PCM 和 WAV 输出为 24 kHz,而 μ-law 和 A-law 输出为 8 kHz,无论响应 mimeType 中的速率是多少。如果流水线需要不同的采样率,请在客户端上重新采样输出。

以下示例请求电话流水线的 μ-law 输出。Google Gen AI SDK for Python 不会在 GenerateContentConfig 中公开 responseFormat,因此 Python 示例会通过 http_options.extra_body 在请求正文中传递它。

Python

from google import genai

client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")

response = client.models.generate_content(
    model="gemini-3.8-flash-tts",
    contents=[{"role": "user", "parts": [{"text": "Have a wonderful day!"}]}],
    config={
        "response_modalities": ["AUDIO"],
        "speech_config": {"voice_config": {"voice": "Kore"}},
        "http_options": {
            "extra_body": {
                "generationConfig": {
                    "responseFormat": [{"audio": {"mimeType": "AUDIO_MULAW"}}]
                }
            }
        },
    },
)

# The response is headerless μ-law audio at 8 kHz, mono.
with open("out.mulaw", "wb") as f:
    f.write(response.candidates[0].content.parts[0].inline_data.data)

REST

curl -X POST \
  -H "Authorization: Bearer $(gcloud auth print-access-token)" \
  -H "Content-Type: application/json" \
  https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
  -d '{
    "contents": [{"role": "user", "parts": [{"text": "Have a wonderful day!"}]}],
    "generationConfig": {
      "responseModalities": ["AUDIO"],
      "responseFormat": [{"audio": {"mimeType": "AUDIO_MULAW"}}],
      "speechConfig": {
        "voiceConfig": {"voice": "Kore"}
      }
    }
  }' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > out.mulaw

支持的语言

Gemini 3.8 TTS 模型会自动检测输入语言。Gemini 3.8 Flash TTS 支持 130 种语言,而 Gemini 3.8 Flash-Lite TTS 支持 101 种语言:

限制

  • TTS 模型接受纯文本输入,并返回纯音频输出。
  • 多音箱请求需要两个音箱。如需在多角色对话中组合设计 (voice_...) 或复制(voice_... 或 voicekey_...)的语音,或使用两个以上的说话者,请单独合成每个说话者的发言,然后将音频串联起来。为每个轮次请求 AUDIO_L16 输出,或者在连接剪辑之前从每个剪辑中剥离 WAV 标头。
  • 这些模型仅在 global 位置提供。
  • 不支持系统指令。
  • 您无法设置输出采样率。不支持 MP3 和 Ogg Opus 输出。

后续步骤