Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, 2 new text-to-speech models in its Gemini Audio family. Google calls them its most expressive audio generation models yet. Flash TTS targets creative direction and character voices. Flash-Lite TTS targets high-volume, cost-efficient production. Both let developers direct delivery line by line using natural language.
Is it deployable? Yes, both models are rolling out now through the Gemini API and Google AI Studio. Access is API-only, with no open weights for self-hosting. Enterprise API access via Gemini Enterprise is listed as coming soon.
What Google Shipped
The release splits TTS into 2 tiers with shared direction controls:
- Gemini 3.8 Flash TTS is built for deep creative direction and character design. Target uses include gaming, immersive audiobooks, podcasts and interactive media. It offers granular control over acting cues, pacing, dialect shifts and backchanneling.
- Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient scale. Google positions it for dubbing, audio content creation and expressive voice agents. It offers fine-grained control over tone, pacing and expressive nuance.
In AI Studio, the playground links use the model identifiers gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.
Voice Design From a Text Prompt
Previous Gemini TTS offered 30 original voices. The 3.8 release moves to a much larger voice system.
- Generative voice design: Flash TTS creates new voices from prompts describing role, accent and voice characteristics. This works across more than 100 languages and dialects. Google’s demos include a Melbourne DJ, a monotone robot and a Japanese dragon.
- Voice library: Developers get 2,000+ production-ready voices. Coverage includes regional varieties like Mexican Spanish, Quebec French and Scots English.
- Save and scale: Custom voices can be saved and reused, with minimal drift across projects.
- Voice remixing (coming soon): Users will adjust a library voice’s timbre, pitch, pace and accent through prompts.
Directing the Performance
Both models accept stage directions written in the script. Gemini can also steer delivery from natural script cues.
- Long-form generation: Voice quality, pacing and timbre hold across hours of continuous audio.
- Native 2-speaker staging: A single script drives a multi-turn conversation with distinct, separated voices.
- Vocal bursts: Non-verbal cues like
<laughs>,<sigh>and<gasp>add conversational texture. - Backchanneling: Active-listening interjections like
|mhm|and|yeah|control reaction beats and comedic timing.
Voice Replication and Safety Controls
Voice replication builds a consistent vocal profile from a 30-second audio sample. The sample must be your voice or one you have rights to use. Replication requires a verbal consent recording from the voice owner, matched against the reference speaker.
Every clip from Gemini Audio models carries a SynthID watermark. This imperceptible mark is embedded directly in the audio output. Replicated voices also carry C2PA content credentials. Google points to the Gemini 3.8 Audio model card for its broader safety approach.
Benchmark Results
Google reports these results for the new models:
- Hume AI Voice Design Benchmark: Flash TTS ranks #1 overall with a score of 71.4, per Hume AI.
- Accent modeling: Flash TTS leads with a score of 60.8.
- Hume AI Overall Quality Index: Flash TTS ranks #1 and Flash-Lite TTS ranks #2.
- Voice Arena blind preference: Both models take top positions in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.
<div id="mtp-g38tts">
<div class="glow"></div>
<h2>How <span class="grad">Gemini 3.8 Flash TTS</span> turns text into a directed performance</h2>
<p class="sub">6 steps, from prompting a new voice to shipping watermarked audio. Waveforms are illustrative visuals, not generated audio.</p>
<div class="steps" role="tablist" aria-label="Explainer steps">
<button class="step on" data-go="0" role="tab">1. Design<b>Prompt a voice</b></button>
<button class="step" data-go="1" role="tab">2. Direct<b>Cue each line</b></button>
<button class="step" data-go="2" role="tab">3. Stage<b>2-speaker scene</b></button>
<button class="step" data-go="3" role="tab">4. Pick<b>Flash or Flash-Lite</b></button>
<button class="step" data-go="4" role="tab">5. Scores<b>Reported results</b></button>
<button class="step" data-go="5" role="tab">6. Protect<b>Consent and SynthID</b></button>
</div>
<section class="slide on" data-i="0">
<h3>Describe a voice in plain language</h3>
<p class="lede">Flash TTS builds a new voice from role, accent and character traits, across more than 100 languages and dialects. Pick one of Google’s own demo prompts.</p>
<div class="stage">
<canvas id="mtp-wave1" aria-hidden="true"></canvas>
<div class="chips" id="mtp-personas">
<button class="chip on" data-p="0">Melbourne DJ</button>
<button class="chip" data-p="1">Tinny robot</button>
<button class="chip" data-p="2">Japanese dragon</button>
</div>
<div class="prompt" id="mtp-prompt"></div>
<div class="spec">
<div><span>Role</span><strong id="mtp-role"></strong></div>
<div><span>Accent or language</span><strong id="mtp-accent"></strong></div>
<div><span>Delivery</span><strong id="mtp-char"></strong></div>
</div>
</div>
<p class="note">Or skip design and pick from 2,000+ production-ready voices, up from 30 originals. Voice remixing (tune timbre, pitch, pace, accent) is coming soon.</p>
</section>
<section class="slide" data-i="1">
<h3>Direct the performance line by line</h3>
<p class="lede">Add cues to the script. Vocal bursts use angle brackets, backchannel interjections use pipes. Watch the waveform react.</p>
<div class="stage">
<canvas id="mtp-wave2" aria-hidden="true"></canvas>
<div class="script" id="mtp-script"></div>
<div class="chips">
<button class="chip" data-cue="d">Stage direction</button>
<button class="chip" data-cue="<laughs>"><laughs></button>
<button class="chip" data-cue="<sigh>"><sigh></button>
<button class="chip" data-cue="<gasp>"><gasp></button>
<button class="chip" data-cue="|mhm|">|mhm|</button>
<button class="chip" data-cue="|yeah|">|yeah|</button>
</div>
<div class="row"><button class="ghost" id="mtp-reset">Clear cues</button></div>
<div class="legend"><span><i style="background:var(–violet)"></i>Stage direction</span><span><i style="background:var(–rose)"></i>Vocal burst</span><span><i style="background:var(–blue)"></i>Backchannel</span></div>
</div>
</section>
<section class="slide" data-i="2">
<h3>Stage a 2-speaker scene from 1 script</h3>
<p class="lede">Both voices stay distinct, turns hand off naturally, and the listener can backchannel while the other speaker talks. Long-form output holds timbre across hours.</p>
<div class="stage">
<div class="lane"><label style="color:var(–blue)">Host</label><div class="track" id="mtp-la"><div class="playhead"></div></div></div>
<div class="lane"><label style="color:var(–violet)">Guest</label><div class="track" id="mtp-lb"><div class="playhead"></div></div></div>
<div class="cap" id="mtp-cap">Press play to run the scene.</div>
<div class="row" style="margin-top:10px"><button class="btn" id="mtp-play">Play scene</button></div>
</div>
</section>
<section class="slide" data-i="3">
<h3>Choose the model for the job</h3>
<p class="lede">Both models share line-by-line direction. They split on what they are tuned for.</p>
<div class="toggle" role="tablist"><button class="on" data-m="0">3.8 Flash TTS</button><button data-m="1">3.8 Flash-Lite TTS</button></div>
<div class="cmp" id="mtp-cmp"></div>
</section>
<section class="slide" data-i="4">
<h3>What Google reports on external leaderboards</h3>
<p class="lede">Scores below are Google-reported results for Gemini 3.8 Flash TTS on Hume AI’s benchmarks and Voice Arena blind preference tests.</p>
<div class="scores">
<div class="score"><div class="n grad" data-to="71.4">0</div><div class="l">Hume AI Voice Design Benchmark, #1 overall</div></div>
<div class="score"><div class="n grad" data-to="60.8">0</div><div class="l">Accent modeling on the same benchmark, leading score</div></div>
</div>
<div class="ranks">
<div class="rank"><div class="medal" style="background:var(–blue)">#1</div><div><strong>3.8 Flash TTS</strong><br><span style="font-size:13px;color:var(–muted)">Hume AI Overall Quality Index</span></div></div>
<div class="rank"><div class="medal" style="background:var(–violet)">#2</div><div><strong>3.8 Flash-Lite TTS</strong><br><span style="font-size:13px;color:var(–muted)">Hume AI Overall Quality Index</span></div></div>
</div>
<p class="note" style="margin-top:14px">Top positions on Voice Arena in:</p>
<div class="langs" id="mtp-langs"><span>Japanese</span><span>Brazilian Portuguese</span><span>Vietnamese</span><span>Modern Standard Arabic</span><span>Mexican Spanish</span><span>Hindi</span></div>
</section>
<section class="slide" data-i="5">
<h3>Replicate a voice, with consent built in</h3>
<p class="lede">Voice replication needs a short sample and a matching consent recording. Every output carries provenance marks.</p>
<div class="flow" id="mtp-flow">
<div class="node"><div class="k">Input</div><strong>30-second sample</strong><p>Your voice, or one you have rights to use.</p></div>
<div class="node"><div class="k">Check</div><strong>Consent recording</strong><p>Verbal consent from the owner must match the reference speaker.</p></div>
<div class="node"><div class="k">Create</div><strong>Saved voice</strong><p>Consistent profile, minimal drift across projects.</p></div>
<div class="node"><div class="k">Output</div><strong>SynthID + C2PA</strong><p>Imperceptible watermark in every clip, plus content credentials.</p></div>
</div>
<div class="row" style="margin-top:12px"><button class="btn" id="mtp-runflow">Run the flow</button></div>
<p class="note">In AI Studio, voice replication is not available in Illinois, Texas, the EEA, the UK, Switzerland and India.</p>
</section>
<div class="nav">
<button class="ghost" id="mtp-prev" disabled>Back</button>
<span id="mtp-count" style="font-size:13px;color:var(–muted)">1 of 6</span>
<button class="btn" id="mtp-next">Next step</button>
</div>
<div class="foot">
<span>Source: <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/" target="_blank" rel="noopener">Google, Sep 23, 2026</a></span>
<a class="brand" href="https://www.marktechpost.com" target="_blank" rel="noopener">Built by Marktechpost</a>
</div>
</div>
<script>
(function(){
var root=document.getElementById(‘mtp-g38tts’);
var reduce=window.matchMedia&&window.matchMedia(‘(prefers-reduced-motion: reduce)’).matches;
function $(s){return root.querySelector(s)}
function $$(s){return Array.prototype.slice.call(root.querySelectorAll(s))}
function sendH(){try{parent.postMessage({type:’mtp-g38tts-height’,height:root.offsetHeight+40},’*’)}catch(e){}}
/* waveform engine */
function Wave(cv){
this.cv=cv;this.ctx=cv.getContext(‘2d’);this.t=0;
this.p={amp:.5,freq:3,jit:.2,hue:0};this.target=Object.assign({},this.p);this.burst=0;this.burstColor=’#D96570′;
var self=this;function loop(){self.draw();requestAnimationFrame(loop)}loop();
}
Wave.prototype.set=function(o){Object.assign(this.target,o)};
Wave.prototype.pulse=function(c){this.burst=1;this.burstColor=c};
Wave.prototype.draw=function(){
var cv=this.cv,ctx=this.ctx,dpr=window.devicePixelRatio||1,w=cv.clientWidth,h=cv.clientHeight;
if(!w)return;
if(cv.width!==w*dpr){cv.width=w*dpr;cv.height=h*dpr}
ctx.setTransform(dpr,0,0,dpr,0,0);ctx.clearRect(0,0,w,h);
for(var k in this.p){this.p[k]+=(this.target[k]-this.p[k])*.06}
this.t+=reduce?0:.035;this.burst*=.965;
var bars=Math.floor(w/7),mid=h/2;
var g=ctx.createLinearGradient(0,0,w,0);g.addColorStop(0,’#4285F4′);g.addColorStop(.55,’#9B72CB’);g.addColorStop(1,’#D96570′);
for(var i=0;i<bars;i++){
var x=i*7+3,u=i/bars;
var env=Math.sin(u*Math.PI);
var v=Math.abs(Math.sin(u*this.p.freq*6+this.t*2)*.6+Math.sin(u*this.p.freq*13-this.t*3)*.4);
v=v*(1-this.p.jit)+this.p.jit*Math.abs(Math.sin(i*12.9898+this.t*4));
var b=this.burst*Math.exp(-Math.pow((u-.5)*5,2));
var bh=Math.max(2,(v*this.p.amp*env+b*.9)*(h*.9));
ctx.fillStyle=b>.15?this.burstColor:g;
ctx.fillRect(x,mid-bh/2,4,bh);
}
};
var w1=new Wave($(‘#mtp-wave1’)),w2=new Wave($(‘#mtp-wave2′));
/* step 1 personas */
var personas=[
{prompt:’A high-energy radio DJ from Melbourne, hyping up the next track.’,role:’Radio DJ’,accent:’Australian English’,ch:’High energy, fast pace’,w:{amp:.95,freq:4.5,jit:.35}},
{prompt:’A super-tinny, monotone robot reading out status updates.’,role:’Robot’,accent:’Synthetic, flat pitch’,ch:’Monotone, tinny timbre’,w:{amp:.4,freq:9,jit:.02}},
{prompt:’A dramatic Japanese dragon speaking from its mountain lair.’,role:’Fire-breathing dragon’,accent:’Japanese’,ch:’Deep, slow, dramatic’,w:{amp:.8,freq:1.4,jit:.12}}
];
var typer;
function persona(i){
var p=personas[i];$$(‘#mtp-personas .chip’).forEach(function(c,j){c.classList.toggle(‘on’,j===i)});
$(‘#mtp-role’).textContent=p.role;$(‘#mtp-accent’).textContent=p.accent;$(‘#mtp-char’).textContent=p.ch;
w1.set(p.w);clearInterval(typer);var el=$(‘#mtp-prompt’),n=0;
if(reduce){el.textContent=p.prompt;return}
typer=setInterval(function(){n++;el.innerHTML=p.prompt.slice(0,n).replace(/</g,’<’)+'<span class="caret"></span>’;if(n>=p.prompt.length)clearInterval(typer)},22);
}
$$(‘#mtp-personas .chip’).forEach(function(c){c.addEventListener(‘click’,function(){persona(+c.dataset.p)})});
persona(0);
/* step 2 script */
var baseWords=[‘I’,’checked’,’the’,’numbers’,’twice,’,’and’,’the’,’launch’,’still’,’goes’,’out’,’on’,’Friday.’];
var cues=[];
function renderScript(){
var html=”,pos={};cues.forEach(function(c){(pos[c.at]=pos[c.at]||[]).push(c)});
for(var i=0;i<=baseWords.length;i++){
(pos[i]||[]).forEach(function(c){html+='<span class="tag ‘+c.k+’">’+c.label.replace(/</g,’<’).replace(/>/g,’>’)+'</span> ‘});
if(i<baseWords.length)html+=baseWords[i]+’ ‘;
}
$(‘#mtp-script’).innerHTML=html;sendH();
}
var slots=[4,13,0,8,6,11];
$$(‘[data-cue]’).forEach(function(b){b.addEventListener(‘click’,function(){
var raw=b.getAttribute(‘data-cue’),k,label,col;
if(raw===’d’){k=’d’;label=’Say this calmly, like a reassuring support agent’;col=’#9B72CB’;w2.set({amp:.45,freq:2,jit:.08})}
else if(raw.charAt(0)===’|’){k=’b’;label=raw;col=’#4285F4′}
else{k=’v’;label=raw;col=’#D96570′}
var at=raw===’d’?0:slots[cues.length%slots.length];
if(raw===’d’&&cues.some(function(c){return c.k===’d’}))return;
cues.push({k:k,label:label,at:at});w2.pulse(col);renderScript();
})});
$(‘#mtp-reset’).addEventListener(‘click’,function(){cues=[];w2.set({amp:.7,freq:3.2,jit:.2});renderScript()});
w2.set({amp:.7,freq:3.2,jit:.2});renderScript();
/* step 3 scene */
var scene=[
{lane:’a’,s:0,e:28,t:’What changed?’,cap:’Host asks the opening question.’},
{lane:’bc’,s:12,e:24,t:’|mhm|’,cap:’Guest backchannels while the host is still talking.’,tr:’b’},
{lane:’b’,s:30,e:64,t:’2 new TTS models’,cap:’Clean handoff to the guest.’},
{lane:’bc’,s:42,e:54,t:’|yeah|’,cap:’Host reacts mid-answer without interrupting.’,tr:’a’},
{lane:’a’,s:66,e:100,t:'<laughs> Nice’,cap:’Scripted laugh lands on cue.’}
];
var playing=false;
$(‘#mtp-play’).addEventListener(‘click’,function(){
if(playing)return;playing=true;var A=$(‘#mtp-la’),B=$(‘#mtp-lb’);
$$(‘#mtp-g38tts .seg’).forEach(function(s){s.remove()});
scene.forEach(function(sg,i){
var tr=sg.lane===’a’?A:sg.lane===’b’?B:(sg.tr===’a’?A:B);
var d=document.createElement(‘div’);d.className=’seg ‘+sg.lane;d.style.left=sg.s+’%’;d.style.width=(sg.e-sg.s)+’%’;d.textContent=sg.t;tr.appendChild(d);
setTimeout(function(){d.classList.add(‘show’);$(‘#mtp-cap’).textContent=sg.cap},reduce?0:i*1100+100);
});
$$(‘#mtp-g38tts .playhead’).forEach(function(p){p.style.opacity=1;p.style.transition=’none’;p.style.left=’0′;void p.offsetWidth;p.style.transition=reduce?’none’:’left 5.6s linear’;p.style.left=’100%’});
setTimeout(function(){playing=false;$(‘#mtp-play’).textContent=’Replay scene’;$$(‘#mtp-g38tts .playhead’).forEach(function(p){p.style.opacity=0})},reduce?50:5800);
});
/* step 4 compare */
var models=[
[[‘Built for’,’Deep creative direction and character design’],[‘Best fit’,’Gaming, immersive audiobooks, podcasts, interactive media’],[‘Control’,’Acting cues, pacing, dialect shifts, backchanneling’],[‘Headline feature’,’Generative voice design from scratch with prompts’],[‘Hume AI Overall Quality Index’,’#1′],[‘Available today’,’Gemini API, Google AI Studio, Gemini Notebook’]],
[[‘Built for’,’High-volume, cost-efficient scale’],[‘Best fit’,’Dubbing, audio content creation, expressive voice agents’],[‘Control’,’Tone, pacing, expressive nuance’],[‘Headline feature’,’Expressive output tuned for volume and cost’],[‘Hume AI Overall Quality Index’,’#2′],[‘Available today’,’Gemini API, Google AI Studio, Google Vids’]]
];
function model(i){
$$(‘.toggle button’).forEach(function(b,j){b.classList.toggle(‘on’,j===i)});
$(‘#mtp-cmp’).innerHTML=models[i].map(function(r,j){return ‘<div style="animation-delay:’+(j*60)+’ms"><span>’+r[0]+'</span><p>’+r[1]+'</p></div>’}).join(”);sendH();
}
$$(‘.toggle button’).forEach(function(b){b.addEventListener(‘click’,function(){model(+b.dataset.m)})});
model(0);
/* step 5 scores */
function runScores(){
$$(‘#mtp-g38tts .score .n’).forEach(function(el){
var to=parseFloat(el.dataset.to),st=null;
if(reduce){el.textContent=to.toFixed(1);return}
function f(ts){if(!st)st=ts;var k=Math.min(1,(ts-st)/1100);el.textContent=(to*(1-Math.pow(1-k,3))).toFixed(1);if(k<1)requestAnimationFrame(f)}
requestAnimationFrame(f);
});
$$(‘#mtp-langs span’).forEach(function(s,i){s.classList.remove(‘show’);setTimeout(function(){s.classList.add(‘show’)},reduce?0:300+i*140)});
}
/* step 6 flow */
function runFlow(){
var n=$$(‘#mtp-flow .node’);n.forEach(function(x){x.classList.remove(‘lit’)});
n.forEach(function(x,i){setTimeout(function(){x.classList.add(‘lit’)},reduce?0:i*650+100)});
}
$(‘#mtp-runflow’).addEventListener(‘click’,runFlow);
/* navigation */
var cur=0,slides=$$(‘.slide’),steps=$$(‘.step’);
function go(i){
cur=Math.max(0,Math.min(slides.length-1,i));
slides.forEach(function(s,j){s.classList.toggle(‘on’,j===cur)});
steps.forEach(function(s,j){s.classList.toggle(‘on’,j===cur);s.classList.toggle(‘done’,j<cur);s.setAttribute(‘aria-selected’,j===cur)});
$(‘#mtp-prev’).disabled=cur===0;$(‘#mtp-next’).disabled=cur===slides.length-1;
$(‘#mtp-count’).textContent=(cur+1)+’ of ‘+slides.length;
if(cur===4)runScores();if(cur===5)runFlow();
setTimeout(sendH,60);
}
steps.forEach(function(s){s.addEventListener(‘click’,function(){go(+s.dataset.go)})});
$(‘#mtp-prev’).addEventListener(‘click’,function(){go(cur-1)});
$(‘#mtp-next’).addEventListener(‘click’,function(){go(cur+1)});
root.addEventListener(‘keydown’,function(e){if(e.target.tagName===’INPUT’)return;if(e.key===’ArrowRight’)go(cur+1);if(e.key===’ArrowLeft’)go(cur-1)});
window.addEventListener(‘resize’,sendH);window.addEventListener(‘load’,sendH);setTimeout(sendH,300);
})();
</script>
“>
Key Takeaways
- Google launched Gemini 3.8 Flash TTS for creative work and Flash-Lite TTS for scale.
- Flash TTS designs new voices from prompts across 100+ languages and dialects.
- Developers get 2,000+ production voices, up from 30 originals.
- Voice replication needs a 30-second sample plus a matching consent recording.
- Flash TTS ranks #1 on Hume AI’s Voice Design Benchmark with 71.4.
FAQ
- What is Gemini 3.8 Flash TTS? It is Google’s text-to-speech model for creative voice design and line-by-line performance direction. It is available through the Gemini API and Google AI Studio.
- How is Flash-Lite TTS different? Flash-Lite TTS is optimized for high-volume, cost-efficient workloads like dubbing and voice agents. It ranks #2 on Hume AI’s Overall Quality Index.
- Can I clone my own voice? Yes, with a 30-second sample and a verbal consent recording. It is unavailable in AI Studio in several regions, including the UK, EEA and India.
Check out the Technical Blog. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design appeared first on MarkTechPost.