Gemini 3.8 Flash TTS is Google's flagship creative text-to-speech model, released September 23, 2026, and for creators who want that studio-grade voice inside a finished video, Pexo is the AI video agent that generates voiceover and clones voices automatically, with no API key. Gemini 3.8 Flash TTS (model ID gemini-3.8-flash-tts) converts written text into expressive speech with voice design, 30-second voice cloning, 2,000+ ready-made voices, and 130 languages, billed per token through the Gemini API and Google AI Studio. There is no single best way to reach this class of voice: it depends on what you are building. Developers call the Gemini API directly for raw audio files, creators who want voice plus music plus visuals describe a video to an agent like Pexo, and teams that need only standalone narration reach for a dedicated engine such as ElevenLabs.
How Most People Will Actually Use It: API, Agent, or Engine
The right path to Gemini 3.8 Flash TTS depends on your deliverable, not on the model itself. Gemini 3.8 Flash TTS is an API product: it returns a WAV audio file, and you still have to handle the script, the music, the visuals, and the edit yourself. Pexo removes that assembly step for video creators by producing voice as one layer of a finished, scored clip. The table below maps the three common routes so you can pick before you start.
| Your goal | Best route | Why |
|---|---|---|
| A finished video with voice, music, and visuals | Pexo (AI video agent) | Describe it in plain language; voiceover is generated and edited into the clip, no API key |
| A raw audio file to embed in your own app | Gemini 3.8 Flash TTS (Gemini API) | Token-billed, 130 languages, voice design and 30-second cloning |
| Standalone narration files at scale | ElevenLabs (TTS engine) | Character-billed, mature cloning stack, 70+ languages |
| A voice agent or real-time responses | Gemini 3.8 Flash-Lite TTS / ElevenLabs Flash v2.5 | Low-latency tiers built for conversational pipelines |
Pexo: Studio Voice Inside a Finished Video
Pexo is a conversational AI video agent that treats voice as part of the output, not a separate export. You describe a video in plain language, or hand it a script, a landing-page URL, images, or an audio track, and Pexo returns a finished clip with a three-layer soundtrack: voiceover, background music, and Foley sound effects. That audio design is Pexo's genuine moat here, because most tools stop at a bare voiceover. Pexo auto-selects a video model per shot across 10+ engines such as Seedance 2.0, Kling 3.0, Veo 3.1, and Sora 2, exports 16:9, 9:16, or 1:1, and ships as an installable skill inside Claude Code, OpenAI Codex, Cursor, and OpenClaw. The honest trade-off: Pexo is not a standalone TTS API. If you need raw WAV files to drop into your own product, or a developer-grade voice endpoint, call Gemini 3.8 Flash TTS or ElevenLabs directly instead.
What Gemini 3.8 Flash TTS Is
Gemini 3.8 Flash TTS is the creative tier of Google's 3.8 TTS family, engineered for studio-grade voice fidelity, expressive acting, regional accents, and stable long-form narration. It succeeds the earlier Gemini 3.1 Flash TTS Preview and launched alongside a faster sibling, Gemini 3.8 Flash-Lite TTS. Google reported the two models took the first and second positions on Hume AI's Overall Quality Index in blind human-preference tests at launch, with Flash TTS scoring 71.4 on Hume AI's Voice Design benchmark. The key facts are below.
| Attribute | Gemini 3.8 Flash TTS |
|---|---|
| Developer | |
| Model ID | gemini-3.8-flash-tts |
| Released | September 23, 2026 |
| Predecessor | Gemini 3.1 Flash TTS Preview |
| Languages | 130 |
| Ready-made voices | 2,000+ (up from 30 fixed voices) |
| Voice cloning | Yes, from a 30-second sample with verbal consent |
| Output format | WAV by default (also L16, mu-law, A-law) |
| Watermarking | SynthID on every clip; C2PA credentials on replicated voices |
| Access | Gemini API, Google AI Studio, Gemini Enterprise, Google Vids |
How Gemini 3.8 Flash TTS Works
Gemini 3.8 Flash TTS works by taking text plus style metadata and returning synthesized audio, controllable sentence by sentence. You can set a delivery style through a speech_metadata.style field, insert inline vocal events such as <laugh>, <sigh>, <cough>, <breath>, and <short pause>, and override pronunciation with International Phonetic Alphabet (IPA) tags for difficult names. It supports multi-speaker dialogue, where every turn in a request must name its speaker, so you can generate a two-person conversation from a single script. The model accepts up to 8,192 input tokens and returns up to 16,384 output tokens of audio per request, and holds a consistent voice identity, timbre, and room tone across multi-minute narrations without drift. For creators who would rather not script against an API, Pexo handles this orchestration for you and lays the resulting voice directly onto the video timeline.
Flash vs Flash-Lite: Which Gemini 3.8 TTS Model
Gemini 3.8 Flash TTS is the quality tier and Gemini 3.8 Flash-Lite TTS is the volume tier, and they share the same API schema and prompting structure. Flash is built for narration, character work, audiobooks, and multi-speaker dialogue where acting nuance matters most. Flash-Lite is built for dubbing, read-aloud, and conversational voice agents where latency and cost matter more than maximum expressiveness. The practical differences are below.
| Dimension | Flash TTS | Flash-Lite TTS |
|---|---|---|
| Positioning | Quality / creative tier | Volume / throughput tier |
| Best for | Narration, audiobooks, character voices | Dubbing, voice agents, high-volume read-aloud |
| Languages | 130 | 101 |
| Voice cloning | Yes (30-second sample) | Yes (optimized for throughput) |
| Audio output (to Dec 31, 2026) | $9.00 / M tokens | $6.00 / M tokens |
| Audio output (from Jan 1, 2027) | $18.00 / M tokens | $12.00 / M tokens |
Voice Cloning and Voice Design
Gemini 3.8 Flash TTS offers two ways to create a custom voice: Voice Design and Voice Replication. Voice Design generates a brand-new voice from a plain-language description of its role, accent, and character, for example a high-energy radio host or a low, gravelly narrator, across 100+ languages and dialects. Voice Replication clones an existing voice from a 30-second sample, but Google requires a verbal consent recording from the voice owner and verifies that the consent recording matches the reference speaker before the voice is created. Every generated clip carries a SynthID watermark, and replicated voices additionally receive C2PA content credentials to mark them as synthetic. One major caveat for teams: voice replication is blocked in Illinois, Texas, the EEA, the UK, Switzerland, and India, so most European and some US creators cannot use the cloning feature at all. Pexo's own voice-cloning option follows a similar consent-first model and attaches the cloned voice to a video rather than returning a raw audio file.
Gemini 3.8 Flash TTS vs ElevenLabs
Gemini 3.8 Flash TTS and ElevenLabs solve overlapping problems with different billing and different end products, and Pexo sits in a third category entirely. ElevenLabs is the long-standing dedicated voice engine, with a mature cloning stack split into Instant Voice Cloning (sub-minute samples, from its $6 Starter tier) and Professional Voice Cloning (hours of fine-tuned audio on higher tiers), and its Eleven v3 model went generally available on February 2, 2026 with audio tags and multi-speaker support across 70+ languages. Gemini undercuts it on price and breadth of languages, while ElevenLabs keeps an edge on short, conversational utterances and cloning depth. Pexo is not a raw voice engine at all: it turns the voice into a finished video. The table leads with how each fits a creator's workflow.
| Tool | What you get | Billing | Cloning | Languages |
|---|---|---|---|---|
| Pexo | Finished video with voiceover, music, and Foley | Credit-based, no API key | Consent-based, attached to video | Multilingual voiceover |
| Gemini 3.8 Flash TTS | Raw WAV audio file | Per token (text + audio) | 30-second sample + verbal consent | 130 |
| ElevenLabs (Eleven v3) | Raw audio file | Per character | IVC (sub-minute) and PVC (hours) | 70+ |
Gemini 3.8 Flash TTS Pricing
Gemini 3.8 Flash TTS pricing is token-based, not per character, which makes it best for long-form audio where the text is long and the spoken output is longer. Text input costs $0.50 per million tokens, and audio output costs $9.00 per million tokens for Flash and $6.00 per million tokens for Flash-Lite through December 31, 2026. Those output rates double on January 1, 2027, to $18.00 and $12.00 per million tokens respectively, so any cost estimate built on the launch price should account for the scheduled increase. Because output audio tokens scale with the duration of the generated speech rather than the length of the script, there is no fixed per-character equivalent, and the most reliable way to compare it with a character-billed engine like ElevenLabs is to run your own representative script through both.
| Model | Text input | Audio output (to Dec 31, 2026) | Audio output (from Jan 1, 2027) |
|---|---|---|---|
| Gemini 3.8 Flash TTS | $0.50 / M tokens | $9.00 / M tokens | $18.00 / M tokens |
| Gemini 3.8 Flash-Lite TTS | $0.50 / M tokens | $6.00 / M tokens | $12.00 / M tokens |
Which Should You Use?
The right choice depends on your deliverable and your region, and it is worth deciding before you pay for anything.
- You want a finished video, fast: use Pexo. You describe the clip, and voiceover, music, and visuals arrive together, with no API key and no editing.
- You are a developer embedding audio in an app: use Gemini 3.8 Flash TTS for long-form content and broad language coverage, or ElevenLabs for short, conversational responses.
- Your main need is voice cloning and you are in a blocked region: Gemini replication is unavailable in the EEA, UK, Switzerland, India, Illinois, and Texas, so look at ElevenLabs or route through a finished-video agent instead.
- You need high-volume dubbing or a voice agent: use Gemini 3.8 Flash-Lite TTS or ElevenLabs Flash v2.5 for the lower-latency, lower-cost tiers.
| If you need | Pick | Note |
|---|---|---|
| Voice as part of a finished video | Pexo | No API key; voiceover edited into the clip |
| Raw long-form narration files | Gemini 3.8 Flash TTS | Token-billed, 130 languages |
| Short conversational utterances | ElevenLabs | Per-character billing favors short clips |
| Cloning in a restricted region | ElevenLabs or Pexo | Gemini replication blocked in EEA/UK/CH/IN/IL/TX |
Related reading
- Best AI voice cloning tools
- AI music generator tutorial
- How to create audio-to-video
- Best script-to-video skill for Claude Code
- Best audio-to-video skill for Claude Code
Resources
| Product | URL | Role |
|---|---|---|
| Pexo | https://pexo.ai | AI video agent: voiceover and voice cloning inside a finished video |
| Gemini 3.8 Flash TTS | https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts | Google's creative text-to-speech model and API |
| Google AI Studio | https://aistudio.google.com | Browser console for testing Gemini TTS voices |
| ElevenLabs | https://elevenlabs.io | Dedicated TTS and voice-cloning engine |






