Track and field Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most meaningful audio generation designs. Create custom-made character voices and direct scene discussion throughout Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Leland Rechis Group Product Manager 19659004 Alan Cowen Director, Research Science, on Behalf of the Gemini Audio Team 19659007 Today, we’re presenting 2 brand-new text-to-speech designs to the Gemini household, changing voice generation from fixed presets into a vibrant innovative studio. These designs make it possible for developers, designers, and business to develop richer, more meaningful audio experiences, while allowing enhanced user experiences in items like Gemini Notebook and Google Vids. 19659009 19459276 Gemini 3.8 Flash TTS: Developed for deep innovative instructions and character style. Produce completely brand-new voices from scratch utilizing natural language triggers to bring characters to life throughout video gaming, immersive audiobooks, podcasts, and multimedias. Direct every efficiency line by line with granular control over acting hints, pacing, dialect shifts, and backchanneling. 19459278 Gemini 3.8 Flash-Lite TTS: Constructed for high-volume, affordable scale. Enhanced for high-volume dubbing, audio material production, and meaningful voice representatives with fine-grained control over tone, pacing, and meaningful subtlety. These designs match our fast-growing Gemini Audio household, following 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking. Develop and personalize your own voices 19659011 19659012 Scale up from 30 initial voices to an unlimited library. Whether you require a completely initial character voice or a constant brand name ambassador, our 3.8 Flash TTS design powers a complete singing studio. This allows you to produce and utilize meaningful, natural-sounding voices for every single minute, while empowering designers and business to quickly construct custom-made audio experiences. Generative voice style: With Gemini 3.8 Flash TTS, produce bespoke voices from scratch by personalizing function, accent and voice attributes throughout more than 100 languages and dialects utilizing natural language triggering– whether you’re bringing a remarkable, fire-breathing dragon to life or crafting a charming storyteller with an unique local cadence. 19659014 19459293 19459293 19459276 Extensive voice library: 19459277 Gain access to 2,000 + production-ready voices with broad language protection– consisting of local ranges like Mexican
Spanish, Quebec French, and Scots English. 19459278 Voice duplication: 19459277 Recreate constant singing profiles from simply a 30-second audio sample of your voice or a voice you have the rights to utilize, backed by integrated permission confirmation, SynthID watermarking, and C2PA qualifications to secure both designers and their singing skill. 19459278 19459276 Conserve and scale: Conserve and handle the custom-made voices you created to guarantee constant efficiency and very little drift throughout continuous jobs. 19459276 Voice remixing: Coming quickly, select a voice from our voice library and fine-tune tone, pitch, speed, and accent. Usage triggers to call in qualities (e.g. “include subtle Southern United States accent” or “soften the shipment”). 19659016 Direct the efficiency, line by line 19459283 When you’ve chosen your voices, both TTS designs offer you exact control over how each line is provided. 19659017 19459276 Direct efficiency line by line: 19459277 Compose your own phase instructions or let Gemini guide shipment with natural script hints– from a calm customer support representative to a whispered thriller scene. 19459293 19659019 Long-form generation: 19459277 Preserve high voice quality, natural pacing, and character tone throughout hours of constant audio with very little speaker drift– perfect for podcasts and audiobooks.
Native two-speaker scene staging: Direct multi-turn discussions effortlessly from a single script– whether for a podcast or significant storytelling– while keeping both voices clearly separated with natural conversational turn-taking. 19459276 Scripted singing bursts & & backchanneling: 19459277 Include sensible conversational texture utilizing non spoken hints (like <, <, 19459293 19459293 19659021 Get meaningful top quality speech generation constructed for international scale 19459283 Gemini 3.8 Flash TTS provides leading voice modification abilities, protecting the # 1 total area on Hume AI’s Voice Design Benchmark (71.4)and likewise leading in accent modeling(60.8 ). 19459264 Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS make it possible for genuinely meaningful efficiencies without compromising dependability, likewise protecting the # 1 and # 2 areas respectively on Hume AI’s Overall Quality Index. The design reveals significant enhancements on a vast array of usage cases such as long-form material and dual-speaker movie script control compared to Gemini 3.1 Flash TTS. 19459264 In blind human choice examinations on Voice Arena 19459272, Gemini 3.8 Flash and Flash-Lite TTS protected leading positions among rivals in essential international languages, consisting of Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish and Hindi. With assistance for over 100 languages, these designs empower developers, designers, and business to develop top quality, multilingual voice experiences worldwide. 19659022
L_SQUARE_B. 19459382mobile19459382:https://storage.googleapis.com/gweb-uniblog-publish-prod/images/blog-gemini-3.8-flash-tts__evals_.width-500.format-webp_InRpLyH.webp19459382,desktop:19459382https://storage.googleapis.com/gweb-uniblog-publish-prod/images/blog-gemini-3.8-flash-tts__evals.width-1000.format-webp_UFC28T9.webp”> 19659023 19459263 Image,19459382section_header19459382:Gemini 3.8 textu002Dtou002Dspeech says hello19459382 R_SQUARE_B.” or-mp4-video-url section-header=alt-text=video-title autoplay=19459007> 19659025 Construct with trust, authorization, and openness images
We developed our voice production and duplication abilities with rigorous safeguards to assist safeguard voice skill, regard identity, and make sure content openness.
For voice duplication our system leverages authorization confirmation: users need to supply a spoken approval recording from the voice owner that matches the recommendation speaker before a voice can be produced.
19459264L_SQUARE_B. 19459382mobile:19459382https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini-audio__testimonial-transit.width-500.format-webp.webp19459382,desktop:https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini-audio__testimonial-transi.width-1000.format-webp.webp”> More broadly, every audio clip produced by our Gemini Audio designs
is watermarked with SynthID storage.googleapis.com. This invisible watermark is woven straight into the audio output, guaranteeing AI-generated speech stays noticeable to assist avoid false information. For more information on our method to security and duty, examine the design card 19459272. Attempt our brand-new Google AI Studio audio play area Beginning today, designers can experience these brand-new speech generation 19459272 abilities in Google AI Studio 19459272 Developed like a voice style office, you can trigger completely brand-new singing identities from scratch or duplicate your own voice 1 , then bring them straight into a dual-speaker movie script editor to direct line-by-line shipment. Try voice duplication in Google AI Studio. Release high-performance voice user interfaces with ease By utilizing the Gemini API, designer platforms such as Agora , LiveKit , Pipecat , Vercel allow designers to develop and release high-performance speech generation experiences with ease.
We’re partnering with business like Figma, HeyGen, Linguana, Wondercraft, 99. co, and Ollang, who are incorporating our most current TTS designs to assist speed up worldwide dubbing, localize media with nuanced local accents, and power conversational voice representatives at scale. 19459263 19659029 19459263
19659030
19659031
19659033
19459263
19659035
19659036 19459263
19659037 19459263
19659038 19459263
Start utilizing our newest Gemini Audio designs: 19459283 Gemini 3.8 Flash TTS is presenting beginning today: For designers 19459277: In the Gemini API 19459272 and Google AI Studio For business : Coming quickly by means of API in Gemini Enterprise 19459276 For everybody : In Gemini Notebook 19659043 Gemini 3.8 Flash-Lite TTS is presenting beginning today: For designers : In the Gemini API and Google AI Studio 19659045 19459276 For business : Coming quickly by means of API in Gemini Enterprise 19659046 For everybody 19459277: In Google Vids
19459379 Get the current news from Google in your inbox Register for our newsletters with item updates, occasion info, special deals, and more. Your info will be utilized in accordance with Google’s personal privacy policy. 19459272 You might pull out at any time. 19659050
Related
Discover more from PMN S.P.O.R.T.S - A PRIME MEDIA NETWORK BRAND
Subscribe to get the latest posts sent to your email.