
DiffusionGemma: The Developer Guide
Google's experimental DiffusionGemma (26B MoE on the Gemma 4 backbone) replaces autoregressive token-by-token generation with diffusion-based parallel decoding on 256-token blocks, delivering up to 4x faster inference, bidirectional context, and built-in self-correction while remaining practical to serve and fine-tune.


