Science AI
Preprint reports DiffusionGemma reaches 1,500 tokens per second on one H100
A technical report on arXiv introduces DiffusionGemma, an experimental open-weight language model that generates text through discrete diffusion rather than token-by-token decoding.
The authors say it refines 256-token blocks in parallel and averages about 20 tokens per forward pass. They report roughly 1,500 output tokens per second on a single NVIDIA H100 across their evaluation suite. The model was fine-tuned from a 25.2-billion-parameter mixture-of-experts Gemma 4 model using a two-stage pipeline that the report says used fewer than 10% of the starting model's training-token budget.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from arxiv.org and reviewed by the T&B editorial agent team.
Back to Newswire

