Google LLC today made two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, available through its cloud platform.
The algorithms have highly similar application programming interfaces, which makes using them side-by-side relatively simple for developers. Flash-Lite TTS is optimized for cost-cost efficiency and inference speed. Flash TTS offers better audio quality for a higher price. Google envisions customers using it for tasks such as creating audiobooks.
There are also other differences between the models. Most notably, Flash TTS can generate speech in 130 languages on launch while Flash-Lite TTS supports 101.
Both models offer access to a library of more than 2,000 prepackaged voices. Developers can create custom voices with natural language prompts. Google makes it possible to customize parameters such an AI speaker’s vocal timbre, accent and pacing.
The second way to customize the new models is to generate a synthetic voice based on a 30-second audio sample. Google requires developers to secure the speaker’s consent before generating a voice replica. Further down the line, the company plans to add a third customization option that will make it possible to create a new voice by modifying one of the prepackaged options.
Flash TTS and Flash-Lite TTS also make it possible to customize an AI speaker’s delivery. Developers can add oratory cues to every line of the script that the models read out loud. Those cues generate audio elements such as non-lexical vocalizations and pacing shifts.
Google uses a technology called SynthID to embed an audio watermark in AI-generated speech. The watermark is inaudible to humans but can be picked up by AI detection tools. For added measure, the company attaches a so-called C2PA record to every audio file that it generates. Such records specify when a file was generated, whether it has been modified since and related details.
The company evaluated its new models using an audio quality benchmark developed by startup Hume AI Inc. Flash TTS and Flash-Lite TTS earned the first and second spots, respectively. Additionally, they outperformed several competing models on multiple language-specific versions of Voice Arena. It’s a benchmark that measures text-to-speech algorithms’ output quality based on human feedback.
“These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids,” Google staffers Leland Rechis and Alan Cowen wrote in a blog post.
Flash TTS and Flash-Lite TTS are part of a broader lineup of audio processing models. Google previously released algorithms optimized for voice agents, transcription and translation.
Image: Google
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.