arXiv:2510.03280cs.LGcs.AI2025-10被引 24
提出首个扩散语言模型的系统性缩放定律,指导训练优化。
Training Optimal Large Diffusion Language Models
- 构建扩散语言模型的计算与数据双约束缩放定律
- 揭示模型规模、数据量与计算量的最优平衡关系
- 为大模型训练提供实用指导,适合研究者参考
我们提出 Quokka,首个针对扩散语言模型(DLMs)的系统性缩放定律,涵盖计算受限与数据受限两种情形,并研究关键建模与优化设计。Quokka 与 Chinchilla 相辅相成,提供更广泛的适用范围。研究成果有望为 DLM 训练提供短期实践指导,并为整个 AI 领域带来长期启发。
原文摘要 · Abstract (English)
We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results would bring short-term practical guidance in DLMs training and long-term inspirations for the whole AI community.
扩散模型语言模型缩放定律
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。