Gemini 3.8 Flash TTS
Give Gemini 3.8 Flash TTS your script and steer mood, pacing and accent line by line — 130 languages, two speakers, right in your browser.
Support
Pro AI Tools
Explore elite tools including the Suno AI Music Generator.
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Gemini 3.8 Flash TTS: A Flagship Tuned for Directed Speech
The September 23, 2026 launch gives Google two speech tiers: a creative flagship and a lean workhorse made for bulk audio jobs.
- A Single Launch, Two Distinct JobsThe expressive flagship handles demanding creative work, while Flash-Lite TTS keeps per-minute costs down for large speech batches.
- Turn-by-Turn Control Over How Lines SoundStyle notes attached to each turn, structured speech metadata and inline vocal events govern tone, pacing, feeling and accent.
- Build a Voice from Text, or Clone One with ConsentSketch a voice with plain-language prompts, or mirror a real speaker using a reference clip alongside a matching consent recording.
Getting Clean Results from Gemini 3.8 Flash TTS
Four habits that keep your transcript tidy while the performance metadata does the heavy lifting.
A Closer Look at Gemini 3.8 Flash TTS Features
Here is what the flagship speech model really delivers, from acting direction to coverage across 130 languages.
Expressive, Line-Level Performance Control
Per-turn styling with inline laughs, sighs, coughs, breaths and pauses feels more like coaching an actor than selecting a preset.
Design Voices with Plain Language
Prompts can specify age, personality, accent, texture and role, with 2,000+ production voices available through the Voices endpoint.
Replication Guarded by Consent
Requires a clean reference clip and a matching consent recording from the same adult speaker, plus SynthID watermarking and C2PA credentials.
Dialogue Scenes Carried by Two Voices
One script can hold a podcast exchange, a teaching dialogue, a product walkthrough or a game scene — no manual line stitching.
Steady Voice Identity Over Long Reads
Google reports consistent identity, timbre, loudness and room tone across multi-minute narration and longer exchanges.
130 Languages, Regional Accents Included
Flash TTS reaches 130 languages versus 101 for Flash-Lite, adding regional accents, minority dialects and IPA pronunciation overrides.
Your Gemini 3.8 Flash TTS Questions, Answered
Quick replies covering cost, model selection, benchmark standings and the safety rules that apply.
What does Gemini 3.8 Flash TTS charge per audio minute?
Launch pricing works out to roughly 1.35 cents per minute of audio, based on $0.50 input and $9 output per million tokens.
Which tier fits my project — Flash TTS or Flash-Lite TTS?
Choose the flagship for nuanced acting and long narration; pick Flash-Lite TTS when you need volume and fast turnaround.
How does it stack up against rival voice models?
Google cites a 71.4 score on Hume's Voice Design Benchmark, while Voice Arena ranks it second with 1,260 Elo.
What changed from the Gemini 3.1 Flash TTS Preview?
Flash-Lite TTS takes over from the 3.1 preview and drops audio output pricing from $20 to $6 per million tokens.
Can I clone a voice, and what safeguards exist?
Cloning requires both a reference clip and a matching consent recording made by the same adult speaker.
Why does the model read my stage directions aloud?
Because the input is treated as a verbatim transcript — shift any lasting directions into the speech metadata field.
Run Your Own Scripts Through Gemini 3.8 Flash TTS
Try either tier inside the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS takes nothing more than swapping a model ID. Weigh batch against priority inference before locking in a production budget.
