Google DeepMind·· 2026-06-11精选AI 评分71
Google DeepMind 发布 DiffusionGemma:文本生成最高提速 4 倍
DiffusionGemma: 4x faster text generation
AI 导读
Google DeepMind 发布实验性开源模型 DiffusionGemma,采用文本扩散方式并行生成整块文本,在专用 GPU 上推理最高提速 4 倍。该模型为 26B MoE 架构,推理时仅激活 3.8B 参数,量化后可放入 18GB 显存,单张 NVIDIA H100 上超过 1000 tokens/秒,RTX 5090 上超过 700 tokens/秒。
推荐理由
官方给出 26B MoE 文本扩散模型的开放权重与硬件实测数据,可据此判断本地低并发推理的取舍。
来源:Google DeepMind · deepmind.google