Gemini 文字转语音 (TTS) 模型可将文字转换为单人或多人音频。您可以使用结构化对话元数据 (speech_metadata) 和内嵌语音标记来控制音频的风格、口音、语速和语气。
本页介绍了如何使用 Gemini Enterprise API 的 generateContent 和 streamGenerateContent 方法通过以下模型生成语音:
- Gemini 3.8 Flash TTS
(
gemini-3.8-flash-tts) - Gemini 3.8 Flash-Lite TTS
(
gemini-3.8-flash-lite-tts)
TTS 模型专为精准朗读文本而量身打造,可对风格和声音进行精细控制,适用于旁白、有声读物和语音代理回答等场景。如需进行交互式实时对话,并使用音频输入和输出,请改用 Live API。
如果您使用的是较早的 Gemini TTS 模型(例如 gemini-3.1-flash-tts-preview),请参阅迁移指南。
支持的模型
| 模型 | 一位说话者 | 多说话人 | 流式 | 语音设计 | 语音复刻 |
|---|---|---|---|---|---|
Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) |
|||||
Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) |
|||||
Gemini 3.1 Flash TTS 预览版 (gemini-3.1-flash-tts-preview) |
|||||
Gemini 2.5 Pro TTS (gemini-2.5-pro-tts) |
|||||
Gemini 2.5 Flash TTS (gemini-2.5-flash-tts) |
本页面上的示例和功能适用于 Gemini 3.8 TTS 模型。对于早期模型,请参阅Google Cloud Text-to-Speech 文档中的 Gemini-TTS。
何时使用哪种模型
Gemini 3.8 TTS 模型具有相同的请求架构和提示格式,因此您可以通过更改模型 ID 在它们之间切换:
- 如果需要优先考虑声音保真度、细致的表演和富有表现力的控制,请使用 Gemini 3.8 Flash TTS (
gemini-3.8-flash-tts)。它适用于工作室级创意工作、复杂的多人对话、频繁的语音片段标记、发音困难的词语、地区或少数民族方言,以及需要稳定声音和室内音调的长篇旁白。 - 使用 Gemini 3.8 Flash-Lite TTS (
gemini-3.8-flash-lite-tts) 可快速高效地处理生产工作负载。它是gemini-3.1-flash-tts-preview的推荐替代方案,适用于大批量批处理制作、对话语音代理级联、朗读功能、语音复制以及主要语言的日常单人语音。
准备工作
- 设置项目并启用 Gemini Enterprise API。
- 为您的开发环境配置应用默认凭证。
如需使用 Python 示例,请安装 Google Gen AI SDK 2.25.0 版或更高版本:
pip install --upgrade "google-genai>=2.25.0"
Gemini 3.8 TTS 模型可在 global 位置使用。向 aiplatform.googleapis.com 端点发送请求。
单说话者 TTS
如需将文本转换为单人音频,请在 parts[].text 中传递逐字转写内容,在 parts[].speech_metadata 中添加可选的轮次级样式,并在 speechConfig.voiceConfig.voice 中设置语音。voice 字段接受预建语音名称、扩展语音库语音 ID,或设计或复制的语音的 ID (voice_...)(或可选的无状态 voicekey_... 键)。
默认情况下,一元请求会返回完整的 WAV 文件(16 位 PCM、24 kHz、单声道),因此您可以将音频字节直接写入 .wav 文件。流式传输请求会返回原始 PCM 块。如需了解详情和其他编码,请参阅音频输出格式。
Python
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
response = client.models.generate_content(
model="gemini-3.8-flash-tts",
contents=[{
"role": "user",
"parts": [{
"text": "Have a wonderful day!",
"speech_metadata": {"style": "cheerful and friendly"},
}],
}],
config={
"response_modalities": ["AUDIO"],
"speech_config": {"voice_config": {"voice": "Kore"}},
},
)
# The SDK has already decoded the base64 audio, so inline_data.data is a
# complete WAV file by default.
with open("out.wav", "wb") as f:
f.write(response.candidates[0].content.parts[0].inline_data.data)
REST
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
-d '{
"contents": [{
"role": "user",
"parts": [{
"text": "Have a wonderful day!",
"speechMetadata": {"style": "cheerful and friendly"}
}]
}],
"generationConfig": {
"responseModalities": ["AUDIO"],
"speechConfig": {
"voiceConfig": {"voice": "Kore"}
}
}
}' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > out.wav
单说话者流式 TTS
如需在模型仍在合成音频时接收音频,请使用 streamGenerateContent 方法。流式传输响应会返回无标头的原始 16 位 PCM 分块(24 kHz,单声道),因此您可以在每个分块到达时将其传递给播放器、套接字或其他音频流水线。在以下 Python 示例中,emit_audio 函数代表该目的地。
Python
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
def emit_audio(pcm: bytes) -> None:
"""Sends a chunk of raw 16-bit PCM audio (24 kHz, mono) downstream."""
# Replace with your audio destination, such as a player, a WebSocket,
# or a telephony stream.
print(f"Received {len(pcm)} bytes of audio")
response_stream = client.models.generate_content_stream(
model="gemini-3.8-flash-lite-tts",
contents=[{
"role": "user",
"parts": [{
"text": "Have a wonderful day!",
"speech_metadata": {"style": "cheerful and friendly"},
}],
}],
config={
"response_modalities": ["AUDIO"],
"speech_config": {"voice_config": {"voice": "Kore"}},
},
)
for chunk in response_stream:
if not chunk.candidates or not chunk.candidates[0].content:
continue
for part in chunk.candidates[0].content.parts or []:
if part.inline_data and part.inline_data.data:
emit_audio(part.inline_data.data)
REST
curl -N -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-lite-tts:streamGenerateContent?alt=sse" \
-d '{
"contents": [{
"role": "user",
"parts": [{
"text": "Have a wonderful day!",
"speechMetadata": {"style": "cheerful and friendly"}
}]
}],
"generationConfig": {
"responseModalities": ["AUDIO"],
"speechConfig": {
"voiceConfig": {"voice": "Kore"}
}
}
}' | sed -n 's/^data: //p' \
| jq -r '.candidates[0].content.parts[0].inlineData.data // empty' \
| while read -r chunk; do echo "$chunk" | base64 --decode; done > streamed.pcm
streamed.pcm 文件包含原始 16 位 PCM 音频,采样率 24 kHz,单声道。
多说话人 TTS
对于两位讲话者之间的对话,请在 speechConfig.multiSpeakerVoiceConfig.speakerVoiceConfigs 中配置两位讲话者。然后,将每个对话轮次作为单独的 part 传递,将每个对话轮次的 speech_metadata.speaker 设置为配置的发言人名称之一,并添加可选的对话轮次级 style。
多音箱请求需要两个音箱。
Python
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
response = client.models.generate_content(
model="gemini-3.8-flash-tts",
contents=[{
"role": "user",
"parts": [
{
"text": "How's it going today, Jane?",
"speech_metadata": {"speaker": "Joe", "style": "cheerful and friendly"},
},
{
"text": "Not too bad, how about you? Ready to test these new voices?",
"speech_metadata": {"speaker": "Jane", "style": "calm and relaxed"},
},
],
}],
config={
"response_modalities": ["AUDIO"],
"speech_config": {
"multi_speaker_voice_config": {
"speaker_voice_configs": [
{"speaker": "Joe", "voice_config": {"voice": "Puck"}},
{"speaker": "Jane", "voice_config": {"voice": "Kore"}},
]
}
},
},
)
with open("dialogue.wav", "wb") as f:
f.write(response.candidates[0].content.parts[0].inline_data.data)
REST
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
-d '{
"contents": [{
"role": "user",
"parts": [
{
"text": "How'\''s it going today, Jane?",
"speechMetadata": {"speaker": "Joe", "style": "cheerful and friendly"}
},
{
"text": "Not too bad, how about you? Ready to test these new voices?",
"speechMetadata": {"speaker": "Jane", "style": "calm and relaxed"}
}
]
}],
"generationConfig": {
"responseModalities": ["AUDIO"],
"speechConfig": {
"multiSpeakerVoiceConfig": {
"speakerVoiceConfigs": [
{"speaker": "Joe", "voiceConfig": {"voice": "Puck"}},
{"speaker": "Jane", "voiceConfig": {"voice": "Kore"}}
]
}
}
}
}' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > dialogue.wav
多说话人流式 TTS
多音箱请求也支持流式传输。使用与多说话人 TTS 中相同的请求,但使用 streamGenerateContent 方法。
Python
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
def emit_audio(pcm: bytes) -> None:
"""Sends a chunk of raw 16-bit PCM audio (24 kHz, mono) downstream."""
# Replace with your audio destination, such as a player, a WebSocket,
# or a telephony stream.
print(f"Received {len(pcm)} bytes of audio")
response_stream = client.models.generate_content_stream(
model="gemini-3.8-flash-tts",
contents=[{
"role": "user",
"parts": [
{
"text": "How's it going today, Jane?",
"speech_metadata": {"speaker": "Joe", "style": "cheerful and friendly"},
},
{
"text": "Not too bad, how about you? Ready to test these new voices?",
"speech_metadata": {"speaker": "Jane", "style": "calm and relaxed"},
},
],
}],
config={
"response_modalities": ["AUDIO"],
"speech_config": {
"multi_speaker_voice_config": {
"speaker_voice_configs": [
{"speaker": "Joe", "voice_config": {"voice": "Puck"}},
{"speaker": "Jane", "voice_config": {"voice": "Kore"}},
]
}
},
},
)
for chunk in response_stream:
if not chunk.candidates or not chunk.candidates[0].content:
continue
for part in chunk.candidates[0].content.parts or []:
if part.inline_data and part.inline_data.data:
emit_audio(part.inline_data.data)
REST
curl -N -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:streamGenerateContent?alt=sse" \
-d '{
"contents": [{
"role": "user",
"parts": [
{
"text": "How'\''s it going today, Jane?",
"speechMetadata": {"speaker": "Joe", "style": "cheerful and friendly"}
},
{
"text": "Not too bad, how about you? Ready to test these new voices?",
"speechMetadata": {"speaker": "Jane", "style": "calm and relaxed"}
}
]
}],
"generationConfig": {
"responseModalities": ["AUDIO"],
"speechConfig": {
"multiSpeakerVoiceConfig": {
"speakerVoiceConfigs": [
{"speaker": "Joe", "voiceConfig": {"voice": "Puck"}},
{"speaker": "Jane", "voiceConfig": {"voice": "Kore"}}
]
}
}
}
}' | sed -n 's/^data: //p' \
| jq -r '.candidates[0].content.parts[0].inlineData.data // empty' \
| while read -r chunk; do echo "$chunk" | base64 --decode; done > dialogue_streamed.pcm
dialogue_streamed.pcm 文件包含原始 16 位 PCM 音频,采样率 24 kHz,单声道。
使用元数据和标记控制语音风格
Gemini 3.8 TTS 模型将 text 字段视为逐字转写内容。写入文本中的说明(例如 "Say cheerfully: Hello!" 或 "Speaker 1: Hello!")可能会被朗读出来。如需控制交付,请按范围拆分说明:
- 持续的轮次级交付 (
speech_metadata.style):将适用于整个轮次的情感、交付风格、韵律、语速和音量放在speech_metadata.style中。例如"style": "whispered urgently"、"style": "out of breath"或"style": "warm and enthusiastic"。 - 时间点事件(内嵌标记):将短暂的非语音声音和停顿放在转写内容中的尖括号内。例如,
"Wait... <short pause> did you hear that? <sigh>"或"Excuse me <cough> as I was saying..."。
如需更多指导,请参阅提示指南。
语音选项
Gemini 3.8 TTS 模型支持四种选择或创建语音的方式:
- 预建语音:30 种精选语音。
- 扩展语音库:包含 2,000 多种精选的预设语音,涵盖多种语言、地区口音和角色人物。
- 语音设计:使用 Voices API
create方法(在 Google Gen AI SDK 中为client.voices.create(),或POST https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices)根据自然语言说明创建自定义语音。将type设置为VOICE_TYPE_PROMPTED,并将store设置为true。响应会返回存储的voice_...ID 和语音样本。 - 语音复刻:使用相同的
create方法,根据参考音频和同意书音频复刻说话人的声音。将type设置为VOICE_TYPE_REPLICATED。使用store: true时,响应会返回存储的voice_...ID。使用store: false时,它会返回无状态的voicekey_...密钥。
自定义语音限制和 TTL
| 语音类型 | 存储模式 | 管理 | 失效时间 |
|---|---|---|---|
存储的语音(voice_...,设计或复制) |
store: true |
Voices API list、get 和 delete 方法 |
自上次使用起 1 年后 |
语音复刻键 (voicekey_...) |
store: false |
由您的应用管理。Google 不会保留密钥。 | 自创建之日起 7 天后 |
使用已存储的声音生成语音会重新开始计算一年的保留期限,因此您经常使用的声音不会过期。语音创建请求受每项目每分钟配额的限制。如果超出此限制,create 方法会返回 RESOURCE_EXHAUSTED 错误。如需了解详情,请参阅配额和系统限制。
预建语音
下表列出了每个预建语音名称及其特征:
- Achernar:柔和
- Achird:友好
- Algenib:Gravelly
- Algieba:平滑
- Alnilam:坚定
- Aoede:轻快
- Autonoe:清亮
- Callirrhoe:随和
- Charon:信息丰富
- Despina:流畅
- Enceladus:气声
- Erinome:清晰
- Fenrir:激昂
- Gacrux:成熟
- Iapetus:清晰
- Kore:坚定
- Laomedeia:欢快
- Leda:青春活力
- Orus:坚固
- Puck:欢快
- Pulcherrima:直率
- Rasalgethi:信息丰富
- Sadachbia:活泼
- Sadaltager:博学
- Schedar:均匀
- Sulafat:暖色调
- Umbriel:轻松
- Vindemiatrix:轻柔
- Zephyr:明快
- Zubenelgenubi:随性
扩展语音库
除了 30 种预建语音外,扩展语音库还提供 2,000 多种精选的预设语音,涵盖多种语言、地区口音、角色人物和使用情形(例如有声读物、对话式代理和新闻)。
如需使用库语音,请在 speechConfig.voiceConfig.voice 中传递其 ID,传递方式与传递预构建语音名称相同。
如需浏览该库,请调用 Voices API list 方法 (ListVoices)。该方法会返回存储在项目中的自定义声音,然后返回预建声音和扩展声音库声音。每个库声音都包含其 id 和元数据,例如 language_code、accent、gender、pitch、persona 和 description。如需缩小结果范围,请使用以下过滤条件:
type:语音来源:prebuilt表示预建语音和库语音,prompted表示设计语音,replicated表示复制语音。 如果您传递多个值,该方法会返回与其中任何一个值匹配的声音。search:要匹配的文本(不区分大小写),用于与每个语音的显示名称和说明进行匹配。
以下示例在库中搜索旁白语音:
Python
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
response = client.voices.list(type_=["prebuilt"], search="narrator", page_size=50)
for voice in response.voices or []:
print(voice.id, voice.language_code, voice.gender, voice.description)
REST
curl -G \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
https://aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/global/voices \
--data-urlencode "type=prebuilt" \
--data-urlencode "search=narrator" \
--data-urlencode "pageSize=50"
list 方法每页最多返回 50 个声音。如需获取下一页,请将响应中的 next_page_token 值作为 page_token 参数传递。
音频输出格式
默认音频编码取决于方法:
- 一元请求 (
generateContent) 会返回包含 RIFF 标头的完整 WAV 文件 (audio/wav):16 位 PCM、24 kHz、单声道。 - 流式请求 (
streamGenerateContent) 返回无标头的原始 16 位有符号小端字节序 PCM 块 (audio/l16),采样率为 24 kHz,单声道。
如需请求其他编码,请将 generationConfig.responseFormat 设置为音频格式。Gemini 3.8 TTS 模型支持以下 mimeType 值:
mimeType 值 |
格式 | 说明 |
|---|---|---|
AUDIO_WAV (一元默认) |
WAV (audio/wav) |
包含 RIFF 标头的完整 WAV 文件(16 位 PCM、24 kHz、单声道)。 |
AUDIO_L16 (流式默认) |
线性 PCM (audio/l16) |
无标头的原始 16 位有符号小端序 PCM。用于流式传输、自定义音频流水线或串联多轮对话片段。 |
AUDIO_MULAW |
μ-law (audio/mulaw) |
G.711 μ-law 压扩音频(8 kHz,单声道),常用于北美和日本的电话。 |
AUDIO_ALAW |
A-law (audio/alaw) |
G.711 A-law 压缩音频(8 kHz,单声道),通常用于欧洲和国际电话。 |
模型会忽略 sampleRate 字段。线性 PCM 和 WAV 输出为 24 kHz,而 μ-law 和 A-law 输出为 8 kHz,无论响应 mimeType 中的速率是多少。如果流水线需要不同的采样率,请在客户端上重新采样输出。
以下示例请求电话流水线的 μ-law 输出。Google Gen AI SDK for Python 不会在 GenerateContentConfig 中公开 responseFormat,因此 Python 示例会通过 http_options.extra_body 在请求正文中传递它。
Python
from google import genai
client = genai.Client(enterprise=True, project="PROJECT_ID", location="global")
response = client.models.generate_content(
model="gemini-3.8-flash-tts",
contents=[{"role": "user", "parts": [{"text": "Have a wonderful day!"}]}],
config={
"response_modalities": ["AUDIO"],
"speech_config": {"voice_config": {"voice": "Kore"}},
"http_options": {
"extra_body": {
"generationConfig": {
"responseFormat": [{"audio": {"mimeType": "AUDIO_MULAW"}}]
}
}
},
},
)
# The response is headerless μ-law audio at 8 kHz, mono.
with open("out.mulaw", "wb") as f:
f.write(response.candidates[0].content.parts[0].inline_data.data)
REST
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/publishers/google/models/gemini-3.8-flash-tts:generateContent \
-d '{
"contents": [{"role": "user", "parts": [{"text": "Have a wonderful day!"}]}],
"generationConfig": {
"responseModalities": ["AUDIO"],
"responseFormat": [{"audio": {"mimeType": "AUDIO_MULAW"}}],
"speechConfig": {
"voiceConfig": {"voice": "Kore"}
}
}
}' | jq -r '.candidates[0].content.parts[0].inlineData.data' | base64 --decode > out.mulaw
支持的语言
Gemini 3.8 TTS 模型会自动检测输入语言。Gemini 3.8 Flash TTS 支持 130 种语言,而 Gemini 3.8 Flash-Lite TTS 支持 101 种语言:
限制
- TTS 模型接受纯文本输入,并返回纯音频输出。
- 多音箱请求需要两个音箱。如需在多角色对话中组合设计 (
voice_...) 或复制(voice_...或voicekey_...)的语音,或使用两个以上的说话者,请单独合成每个说话者的发言,然后将音频串联起来。为每个轮次请求AUDIO_L16输出,或者在连接剪辑之前从每个剪辑中剥离 WAV 标头。 - 这些模型仅在
global位置提供。 - 不支持系统指令。
- 您无法设置输出采样率。不支持 MP3 和 Ogg Opus 输出。
后续步骤
- 如需了解如何指导风格、节奏和声音,请参阅提示指南。
- 借助语音设计功能,根据文字说明创建自定义语音。
- 使用语音复刻功能复刻说话者的声音。
- 如需查看模型规范,请访问 Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 模型页面。