Google DeepMind has transformed its Gemma 4 model into a new text diffusion model named DiffusionGemma, utilizing less than 10 percent of the original training budget, according to The Decoder. This retrofit represents a significant reduction in computational resources compared to the initial training process.

DiffusionGemma is capable of generating 256 tokens simultaneously, rather than sequentially, achieving a throughput of approximately 1,500 tokens per second. Despite this impressive speed, The Decoder notes that the model’s quality still lags behind the original autoregressive Gemma 4, especially in reasoning tasks where accuracy is critical.

For Japanese markets, where AI-driven language models are increasingly applied in financial analysis and automated trading systems, such efficiency improvements could lower costs and accelerate deployment of AI tools across FX, crypto, and equities sectors.