Google has unveiled two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, which support over 100 languages. These models bring enhanced capabilities for voice synthesis, offering features such as voice creation from text descriptions, stage directions, and generating dialogues between two voices.

According to The Decoder, the technology also includes voice cloning from just a 30-second audio sample, enabling more personalized and realistic voice outputs. This advancement could have significant applications in areas like automated customer service, content creation, and accessibility tools.

For Japanese markets, where multilingual communication and voice technology adoption continue to grow, these developments by Google could accelerate integration of sophisticated TTS solutions in FX, crypto, and equity platforms, improving user experience and operational efficiency.