Gemini 3.8 Flash TTS

Give Gemini 3.8 Flash TTS your script and steer mood, pacing and accent line by line — 130 languages, two speakers, right in your browser.

Gemini 3.8 Flash TTS
Craft expressive speech with Flash TTS, or keep costs low at scale with Flash-Lite TTS
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

Gemini 3.8 Flash TTS: A Flagship Tuned for Directed Speech

The September 23, 2026 launch gives Google two speech tiers: a creative flagship and a lean workhorse made for bulk audio jobs.

  • A Single Launch, Two Distinct Jobs
    The expressive flagship handles demanding creative work, while Flash-Lite TTS keeps per-minute costs down for large speech batches.
  • Turn-by-Turn Control Over How Lines Sound
    Style notes attached to each turn, structured speech metadata and inline vocal events govern tone, pacing, feeling and accent.
  • Build a Voice from Text, or Clone One with Consent
    Sketch a voice with plain-language prompts, or mirror a real speaker using a reference clip alongside a matching consent recording.

Getting Clean Results from Gemini 3.8 Flash TTS

Four habits that keep your transcript tidy while the performance metadata does the heavy lifting.

A Closer Look at Gemini 3.8 Flash TTS Features

Here is what the flagship speech model really delivers, from acting direction to coverage across 130 languages.

Expressive, Line-Level Performance Control

Per-turn styling with inline laughs, sighs, coughs, breaths and pauses feels more like coaching an actor than selecting a preset.

Design Voices with Plain Language

Prompts can specify age, personality, accent, texture and role, with 2,000+ production voices available through the Voices endpoint.

Replication Guarded by Consent

Requires a clean reference clip and a matching consent recording from the same adult speaker, plus SynthID watermarking and C2PA credentials.

Dialogue Scenes Carried by Two Voices

One script can hold a podcast exchange, a teaching dialogue, a product walkthrough or a game scene — no manual line stitching.

Steady Voice Identity Over Long Reads

Google reports consistent identity, timbre, loudness and room tone across multi-minute narration and longer exchanges.

130 Languages, Regional Accents Included

Flash TTS reaches 130 languages versus 101 for Flash-Lite, adding regional accents, minority dialects and IPA pronunciation overrides.

FAQ

Your Gemini 3.8 Flash TTS Questions, Answered

Quick replies covering cost, model selection, benchmark standings and the safety rules that apply.

1

What does Gemini 3.8 Flash TTS charge per audio minute?

Launch pricing works out to roughly 1.35 cents per minute of audio, based on $0.50 input and $9 output per million tokens.

2

Which tier fits my project — Flash TTS or Flash-Lite TTS?

Choose the flagship for nuanced acting and long narration; pick Flash-Lite TTS when you need volume and fast turnaround.

3

How does it stack up against rival voice models?

Google cites a 71.4 score on Hume's Voice Design Benchmark, while Voice Arena ranks it second with 1,260 Elo.

4

What changed from the Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS takes over from the 3.1 preview and drops audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what safeguards exist?

Cloning requires both a reference clip and a matching consent recording made by the same adult speaker.

6

Why does the model read my stage directions aloud?

Because the input is treated as a verbatim transcript — shift any lasting directions into the speech metadata field.

Run Your Own Scripts Through Gemini 3.8 Flash TTS

Try either tier inside the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS takes nothing more than swapping a model ID. Weigh batch against priority inference before locking in a production budget.