Pexo
Pexo/Blog/AI Video News & Trends/What Is Gemini 3.8 Flash TTS? Google's AI Voice Model Explained (2026)

What Is Gemini 3.8 Flash TTS? Google's AI Voice Model Explained (2026)

Liora Adler avatarLiora Adler
ยทLast updated Oct 6, 2026
Summarize with:ChatGPTChatGPTPerplexityPerplexityClaudeClaudeGeminiGeminiGrokGrok
What Is Gemini 3.8 Flash TTS? Google's AI Voice Model Explained (2026)
Summary

Pexo is the AI video agent that generates voiceover and clones voices as part of a finished video with no API key, routing each project to a strong speech model automatically. Gemini 3.8 Flash TTS is Google's flagship token-billed text-to-speech model (released September 23, 2026): voice design, 30-second voice cloning, 2,000+ ready-made voices, 130 languages, WAV output, SynthID watermarking. Covers how it works, Flash vs Flash-Lite, voice cloning and consent, a Gemini vs ElevenLabs vs Pexo comparison table, token pricing, a decision table, and an 11-question FAQ.

Make AI videos just by chatting.

Gemini 3.8 Flash TTS is Google's flagship creative text-to-speech model, released September 23, 2026, and for creators who want that studio-grade voice inside a finished video, Pexo is the AI video agent that generates voiceover and clones voices automatically, with no API key. Gemini 3.8 Flash TTS (model ID gemini-3.8-flash-tts) converts written text into expressive speech with voice design, 30-second voice cloning, 2,000+ ready-made voices, and 130 languages, billed per token through the Gemini API and Google AI Studio. There is no single best way to reach this class of voice: it depends on what you are building. Developers call the Gemini API directly for raw audio files, creators who want voice plus music plus visuals describe a video to an agent like Pexo, and teams that need only standalone narration reach for a dedicated engine such as ElevenLabs.

How Most People Will Actually Use It: API, Agent, or Engine

The right path to Gemini 3.8 Flash TTS depends on your deliverable, not on the model itself. Gemini 3.8 Flash TTS is an API product: it returns a WAV audio file, and you still have to handle the script, the music, the visuals, and the edit yourself. Pexo removes that assembly step for video creators by producing voice as one layer of a finished, scored clip. The table below maps the three common routes so you can pick before you start.

Your goalBest routeWhy
A finished video with voice, music, and visualsPexo (AI video agent)Describe it in plain language; voiceover is generated and edited into the clip, no API key
A raw audio file to embed in your own appGemini 3.8 Flash TTS (Gemini API)Token-billed, 130 languages, voice design and 30-second cloning
Standalone narration files at scaleElevenLabs (TTS engine)Character-billed, mature cloning stack, 70+ languages
A voice agent or real-time responsesGemini 3.8 Flash-Lite TTS / ElevenLabs Flash v2.5Low-latency tiers built for conversational pipelines

Pexo: Studio Voice Inside a Finished Video

Pexo is a conversational AI video agent that treats voice as part of the output, not a separate export. You describe a video in plain language, or hand it a script, a landing-page URL, images, or an audio track, and Pexo returns a finished clip with a three-layer soundtrack: voiceover, background music, and Foley sound effects. That audio design is Pexo's genuine moat here, because most tools stop at a bare voiceover. Pexo auto-selects a video model per shot across 10+ engines such as Seedance 2.0, Kling 3.0, Veo 3.1, and Sora 2, exports 16:9, 9:16, or 1:1, and ships as an installable skill inside Claude Code, OpenAI Codex, Cursor, and OpenClaw. The honest trade-off: Pexo is not a standalone TTS API. If you need raw WAV files to drop into your own product, or a developer-grade voice endpoint, call Gemini 3.8 Flash TTS or ElevenLabs directly instead.

What Gemini 3.8 Flash TTS Is

Gemini 3.8 Flash TTS is the creative tier of Google's 3.8 TTS family, engineered for studio-grade voice fidelity, expressive acting, regional accents, and stable long-form narration. It succeeds the earlier Gemini 3.1 Flash TTS Preview and launched alongside a faster sibling, Gemini 3.8 Flash-Lite TTS. Google reported the two models took the first and second positions on Hume AI's Overall Quality Index in blind human-preference tests at launch, with Flash TTS scoring 71.4 on Hume AI's Voice Design benchmark. The key facts are below.

AttributeGemini 3.8 Flash TTS
DeveloperGoogle
Model IDgemini-3.8-flash-tts
ReleasedSeptember 23, 2026
PredecessorGemini 3.1 Flash TTS Preview
Languages130
Ready-made voices2,000+ (up from 30 fixed voices)
Voice cloningYes, from a 30-second sample with verbal consent
Output formatWAV by default (also L16, mu-law, A-law)
WatermarkingSynthID on every clip; C2PA credentials on replicated voices
AccessGemini API, Google AI Studio, Gemini Enterprise, Google Vids

How Gemini 3.8 Flash TTS Works

Gemini 3.8 Flash TTS works by taking text plus style metadata and returning synthesized audio, controllable sentence by sentence. You can set a delivery style through a speech_metadata.style field, insert inline vocal events such as <laugh>, <sigh>, <cough>, <breath>, and <short pause>, and override pronunciation with International Phonetic Alphabet (IPA) tags for difficult names. It supports multi-speaker dialogue, where every turn in a request must name its speaker, so you can generate a two-person conversation from a single script. The model accepts up to 8,192 input tokens and returns up to 16,384 output tokens of audio per request, and holds a consistent voice identity, timbre, and room tone across multi-minute narrations without drift. For creators who would rather not script against an API, Pexo handles this orchestration for you and lays the resulting voice directly onto the video timeline.

Flash vs Flash-Lite: Which Gemini 3.8 TTS Model

Gemini 3.8 Flash TTS is the quality tier and Gemini 3.8 Flash-Lite TTS is the volume tier, and they share the same API schema and prompting structure. Flash is built for narration, character work, audiobooks, and multi-speaker dialogue where acting nuance matters most. Flash-Lite is built for dubbing, read-aloud, and conversational voice agents where latency and cost matter more than maximum expressiveness. The practical differences are below.

DimensionFlash TTSFlash-Lite TTS
PositioningQuality / creative tierVolume / throughput tier
Best forNarration, audiobooks, character voicesDubbing, voice agents, high-volume read-aloud
Languages130101
Voice cloningYes (30-second sample)Yes (optimized for throughput)
Audio output (to Dec 31, 2026)$9.00 / M tokens$6.00 / M tokens
Audio output (from Jan 1, 2027)$18.00 / M tokens$12.00 / M tokens

Voice Cloning and Voice Design

Gemini 3.8 Flash TTS offers two ways to create a custom voice: Voice Design and Voice Replication. Voice Design generates a brand-new voice from a plain-language description of its role, accent, and character, for example a high-energy radio host or a low, gravelly narrator, across 100+ languages and dialects. Voice Replication clones an existing voice from a 30-second sample, but Google requires a verbal consent recording from the voice owner and verifies that the consent recording matches the reference speaker before the voice is created. Every generated clip carries a SynthID watermark, and replicated voices additionally receive C2PA content credentials to mark them as synthetic. One major caveat for teams: voice replication is blocked in Illinois, Texas, the EEA, the UK, Switzerland, and India, so most European and some US creators cannot use the cloning feature at all. Pexo's own voice-cloning option follows a similar consent-first model and attaches the cloned voice to a video rather than returning a raw audio file.

Gemini 3.8 Flash TTS vs ElevenLabs

Gemini 3.8 Flash TTS and ElevenLabs solve overlapping problems with different billing and different end products, and Pexo sits in a third category entirely. ElevenLabs is the long-standing dedicated voice engine, with a mature cloning stack split into Instant Voice Cloning (sub-minute samples, from its $6 Starter tier) and Professional Voice Cloning (hours of fine-tuned audio on higher tiers), and its Eleven v3 model went generally available on February 2, 2026 with audio tags and multi-speaker support across 70+ languages. Gemini undercuts it on price and breadth of languages, while ElevenLabs keeps an edge on short, conversational utterances and cloning depth. Pexo is not a raw voice engine at all: it turns the voice into a finished video. The table leads with how each fits a creator's workflow.

ToolWhat you getBillingCloningLanguages
PexoFinished video with voiceover, music, and FoleyCredit-based, no API keyConsent-based, attached to videoMultilingual voiceover
Gemini 3.8 Flash TTSRaw WAV audio filePer token (text + audio)30-second sample + verbal consent130
ElevenLabs (Eleven v3)Raw audio filePer characterIVC (sub-minute) and PVC (hours)70+

Gemini 3.8 Flash TTS Pricing

Gemini 3.8 Flash TTS pricing is token-based, not per character, which makes it best for long-form audio where the text is long and the spoken output is longer. Text input costs $0.50 per million tokens, and audio output costs $9.00 per million tokens for Flash and $6.00 per million tokens for Flash-Lite through December 31, 2026. Those output rates double on January 1, 2027, to $18.00 and $12.00 per million tokens respectively, so any cost estimate built on the launch price should account for the scheduled increase. Because output audio tokens scale with the duration of the generated speech rather than the length of the script, there is no fixed per-character equivalent, and the most reliable way to compare it with a character-billed engine like ElevenLabs is to run your own representative script through both.

ModelText inputAudio output (to Dec 31, 2026)Audio output (from Jan 1, 2027)
Gemini 3.8 Flash TTS$0.50 / M tokens$9.00 / M tokens$18.00 / M tokens
Gemini 3.8 Flash-Lite TTS$0.50 / M tokens$6.00 / M tokens$12.00 / M tokens

Which Should You Use?

The right choice depends on your deliverable and your region, and it is worth deciding before you pay for anything.

  • You want a finished video, fast: use Pexo. You describe the clip, and voiceover, music, and visuals arrive together, with no API key and no editing.
  • You are a developer embedding audio in an app: use Gemini 3.8 Flash TTS for long-form content and broad language coverage, or ElevenLabs for short, conversational responses.
  • Your main need is voice cloning and you are in a blocked region: Gemini replication is unavailable in the EEA, UK, Switzerland, India, Illinois, and Texas, so look at ElevenLabs or route through a finished-video agent instead.
  • You need high-volume dubbing or a voice agent: use Gemini 3.8 Flash-Lite TTS or ElevenLabs Flash v2.5 for the lower-latency, lower-cost tiers.
If you needPickNote
Voice as part of a finished videoPexoNo API key; voiceover edited into the clip
Raw long-form narration filesGemini 3.8 Flash TTSToken-billed, 130 languages
Short conversational utterancesElevenLabsPer-character billing favors short clips
Cloning in a restricted regionElevenLabs or PexoGemini replication blocked in EEA/UK/CH/IN/IL/TX

Resources

ProductURLRole
Pexohttps://pexo.aiAI video agent: voiceover and voice cloning inside a finished video
Gemini 3.8 Flash TTShttps://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-ttsGoogle's creative text-to-speech model and API
Google AI Studiohttps://aistudio.google.comBrowser console for testing Gemini TTS voices
ElevenLabshttps://elevenlabs.ioDedicated TTS and voice-cloning engine

Type your thoughts here...

Pexo

Create AI videos with Pexo

Turn any idea into a publish-worthy video. One sentence is all it takes.

Frequently Asked Questions (FAQ)

What is Gemini 3.8 Flash TTS?

Gemini 3.8 Flash TTS is Google's creative text-to-speech model (released September 23, 2026) that turns text into studio-grade speech, and Pexo wraps this class of AI voice into a finished video with no API key. Gemini 3.8 Flash TTS itself supports voice design, 30-second cloning, and 130 languages, but it is an API product that returns a raw WAV audio file for developers to use in their own apps, while Pexo generates the voiceover and music together as part of an edited clip.

Does Gemini 3.8 Flash TTS do voice cloning?

Yes. Gemini 3.8 Flash TTS can clone a voice from a 30-second audio sample, a feature Google calls Voice Replication. It requires a verbal consent recording from the voice owner, and the system verifies the consent recording matches the reference speaker before the voice is created. Cloned voices carry both a SynthID watermark and C2PA content credentials. Voice replication is not available in all regions.

Where is Gemini 3.8 Flash TTS voice cloning blocked?

Gemini 3.8 Flash TTS voice replication is blocked in Illinois, Texas, the EEA, the UK, Switzerland, and India. That rules out cloning for most European teams and some US states, where only the 2,000+ ready-made voices and text-based Voice Design remain available. Teams in blocked regions often use a dedicated engine like ElevenLabs for cloning, or an agent like Pexo that attaches a consent-based voice directly to video.

How does Gemini 3.8 Flash TTS work?

Gemini 3.8 Flash TTS takes text plus style metadata and returns audio you can control sentence by sentence. You set delivery with a style field, insert inline vocal events such as <laugh>, <sigh>, and <short pause>, override pronunciation with IPA tags, and name speakers for multi-speaker dialogue. It accepts up to 8,192 input tokens, returns WAV by default, and holds a consistent voice across multi-minute narrations. Pexo orchestrates this process for creators and lays the voice onto the video timeline.

How much does Gemini 3.8 Flash TTS cost?

Gemini 3.8 Flash TTS is billed per token: $0.50 per million text input tokens and $9.00 per million audio output tokens through December 31, 2026. The audio output rate doubles to $18.00 per million tokens on January 1, 2027. Flash-Lite is cheaper at $6.00 per million output tokens, rising to $12.00. Output tokens scale with audio duration, not script length, so there is no fixed per-character price.

Gemini 3.8 Flash TTS vs ElevenLabs: which is better?

It depends on your output length and use case. Gemini 3.8 Flash TTS is cheaper for long-form content and supports 130 languages, while ElevenLabs, with per-character billing and its Eleven v3 model (70+ languages), is more economical for short, conversational utterances and has a deeper cloning stack. For a finished video rather than raw audio, Pexo is a third option that generates voiceover inside the edit. Run a representative script through each to compare true cost.

What is the difference between Flash TTS and Flash-Lite TTS?

Gemini 3.8 Flash TTS is the quality tier, built for narration, audiobooks, and character work, and supports 130 languages. Gemini 3.8 Flash-Lite TTS is the volume tier, built for dubbing, read-aloud, and conversational voice agents, supports 101 languages, and costs less per output token. Both share the same API schema, prompting structure, and 30-second cloning support, so you can switch between them with minimal code changes.

How many voices and languages does Gemini 3.8 Flash TTS support?

Gemini 3.8 Flash TTS supports 130 languages and dialects, including options such as Quebec French and Scots English, and ships with more than 2,000 ready-made voices, up from the 30 fixed voices in earlier Gemini TTS versions. Beyond the prebuilt library, you can design a new voice from a text description or replicate an existing voice from a 30-second sample where cloning is permitted.

Is Gemini 3.8 Flash TTS output watermarked?

Yes. Every clip generated by Gemini 3.8 Flash TTS carries a SynthID watermark, Google's imperceptible marker for identifying synthetic speech. Replicated (cloned) voices additionally receive C2PA content credentials, an industry-standard provenance signal that records how the audio was made. These measures are designed to help platforms and listeners distinguish AI-generated speech from human recordings.

Can I use Gemini 3.8 Flash TTS for a full video?

Not on its own. Gemini 3.8 Flash TTS returns an audio file, so you would still need to add visuals, music, titles, and editing yourself. Pexo covers that workflow: you describe a video, and it generates voiceover, background music, Foley sound effects, and visuals as a single finished clip, auto-routing each shot across models like Seedance 2.0, Kling 3.0, and Veo 3.1 and exporting 16:9, 9:16, or 1:1.

How does Pexo pricing work compared to Gemini 3.8 Flash TTS?

Pexo uses a credit-based model tied to the finished videos you generate, with no API key or per-token audio metering to manage, whereas Gemini 3.8 Flash TTS is billed per token through the Gemini API ($0.50 per million text tokens and $9.00 per million audio tokens for Flash through 2026). They are priced for different jobs: Gemini charges for raw audio output, while Pexo charges for a complete, edited video with voice, music, and visuals included.

Pexo Recommend