Is Gemini 3.8 TTS safe? Voice cloning, consent and SynthID
Google’s Gemini 3.8 Flash TTS can copy a voice from 30 seconds of audio. What stops misuse, where the watermark helps, and what it cannot do for you.
Short answer: Gemini 3.8 TTS is safer than most voice tools to use, and no safer than any of them to be on the receiving end of. Google has put two real controls on it. Every clip carries an inaudible watermark, and copying a real person’s voice needs a spoken consent recording from that person. Neither control helps much when someone plays you a voice over the phone. That part is still on you.
Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026, in its launch post, rolling out the same day in the Gemini API and Google AI Studio. Everything below is as of 25 September 2026.
What the two models do
Both turn text into speech. Flash is aimed at creative work, such as audiobooks, podcasts and game characters, where you direct the performance line by line. Flash-Lite is the cheaper, high-volume version for dubbing and voice agents. Google lists more than 2,000 ready-made voices across over 100 languages and dialects, two-speaker scenes, and scripted sounds like laughs and sighs.
The feature that raises the safety question is voice replication. According to Google, the model can build a consistent copy of a voice from a 30-second audio sample. You can also describe a voice in plain words and have one designed from scratch, which is a different thing: an invented voice belongs to nobody.
The two safeguards Google built in
- Consent recording. Google says a voice can only be replicated after the user provides a spoken consent recording from the voice’s owner, and that recording has to match the reference speaker. A stolen clip of someone’s voice is not enough on its own.
- Watermarking. Google says every audio clip from its Gemini audio models is marked with SynthID, a watermark people cannot hear. Google DeepMind’s SynthID page says the audio mark survives common changes like added noise, MP3 compression and speed changes.
- Regional limits. Voice replication through AI Studio is not available in Illinois, Texas, the European Economic Area, the UK, Switzerland or India. Google does not say why. Several of those places have laws on biometric data, which is a plausible reason, not a stated one.
- Usage rules. Google’s Generative AI Prohibited Use Policy applies to these models, and it bans impersonating people to deceive.
These are sensible choices, and not every voice tool makes them. The consent check in particular puts real friction in front of the most obvious abuse, which is cloning a specific person without asking.
What the model card says, and what it leaves out
The Gemini 3.8 Audio model card covers the TTS models alongside Gemini 3.8 Live. It describes automated and human safety testing, separate child-safety evaluations, and ongoing work on resisting jailbreaks. Under Google’s frontier safety framework, it concludes the TTS models are not likely to reach any of the critical capability levels it tracks.
What the card does not discuss is impersonation. It says nothing specific about voice cloning, fraud or misinformation, and it does not mention SynthID or the consent check. Those appear only in the launch post. That is not proof of a gap in testing. It does mean the most relevant risk for a voice model is documented in marketing copy rather than in the technical safety document.
A watermark proves a clip came from a machine. It does not stop the clip from being played down a phone line.
Where the watermark stops helping
SynthID is useful when someone checks for it. Google says you can upload an audio clip to the Gemini app and ask whether it carries a SynthID mark, and it runs a separate SynthID Detector portal that is still in early testing with journalists. That works well for a newsroom verifying a recording. It does nothing during a live call from a number you don’t recognise.
It also only covers Google’s own tools. A voice cloned with some other service carries no SynthID mark, so a clip that comes back clean has not been proven human. It has only been shown not to come from Google.
The scam that matters here is older than this model. The US Federal Trade Commission warned in 2023 that criminals clone a family member’s voice from clips posted online and phone a relative with a fake emergency. Its advice is blunt: don’t trust the voice, and call the person back on a number you know is theirs. That advice does not depend on which company made the voice. We cover how to talk older relatives through it in AI for seniors.
A worked example: the urgent voice note
A finance assistant gets a voice note that sounds exactly like the managing director: “I’m about to board, I need you to pay the new supplier invoice today, I’ll explain later.” The voice is right. The urgency is plausible. The details come by email a minute later.
- Treat the voice as unverified, however familiar it sounds.
- Call the director back on the number already in the company directory, not the one in the message.
- Check whether the request breaks a normal rule, such as a new supplier with no purchase order. Pressure to skip a step is the warning sign.
- If you want to know where the clip came from, upload it to the Gemini app and ask about SynthID. A match tells you it was made with a Google tool. No match tells you very little.
The habit is the same one you need for text answers: check the claim against a source you trust before acting on it. We walk through that for written output in how to check an AI answer when you are not the expert.
If you want to use it at work
- Prefer designed voices to cloned ones. An invented voice has no owner to harm and no consent to manage.
- If you clone a colleague’s voice for training videos, get their consent in writing as well as the recording Google asks for, and agree where the voice may be used.
- Tell listeners the voice is synthetic. The watermark is invisible to them.
- For customer-facing voice agents, plan what happens when the agent is wrong. The trade-offs are the same ones we cover in generative AI in customer service.
- Remember what AI voices are still bad at: they read what they are given with total confidence. More on that in what AI is actually bad at.
The verdict
For someone using it, Gemini 3.8 TTS is one of the more carefully fenced voice tools on the market: consent before cloning, a watermark on every clip, and regional limits where the law is strictest. For everyone else, the risk from cloned voices is the same as it was last week, because it never depended on one vendor. We read the documents for OpenAI’s latest models the same way in is GPT-6 Sol safe and for Meta’s agent in is Meta Muse safe.
The skill that holds up across every new model is the unexciting one: verify before you act. That is what Coursium teaches, on your phone, one short lesson at a time.
Frequently asked questions
Can Gemini 3.8 TTS clone anyone’s voice?
Not without their participation, according to Google. Voice replication needs a 30-second sample plus a spoken consent recording from the voice’s owner that matches the sample. Replication through AI Studio is also unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India.
How can I tell if audio was made with Gemini?
Google says every clip from its Gemini audio models carries a SynthID watermark. You can upload a clip to the Gemini app and ask it to check. A match means a Google tool made it. No match does not prove a human did, because other tools do not use SynthID.
Does the model card cover voice cloning risks?
No. The Gemini 3.8 Audio model card covers safety testing, child safety and frontier risk, and concludes the models are unlikely to reach critical capability levels. The consent check and SynthID are described in Google’s launch post, not the model card.
What should I do if I get a call in a familiar voice asking for money?
Follow the US FTC’s advice: do not trust the voice, and call the person back on a number you already know is theirs. Requests for wire transfers, gift cards or cryptocurrency are a strong warning sign.