arXiv:2510.03280cs.LGcs.AI2025-10被引 24

提出首个扩散语言模型的系统性缩放定律,指导训练优化。

Training Optimal Large Diffusion Language Models

  • 构建扩散语言模型的计算与数据双约束缩放定律
  • 揭示模型规模、数据量与计算量的最优平衡关系
  • 为大模型训练提供实用指导,适合研究者参考

我们提出 Quokka,首个针对扩散语言模型(DLMs)的系统性缩放定律,涵盖计算受限与数据受限两种情形,并研究关键建模与优化设计。Quokka 与 Chinchilla 相辅相成,提供更广泛的适用范围。研究成果有望为 DLM 训练提供短期实践指导,并为整个 AI 领域带来长期启发。

原文摘要 · Abstract (English)

We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results would bring short-term practical guidance in DLMs training and long-term inspirations for the whole AI community.

扩散模型语言模型缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。