7B参数的统一扩散语言模型Sumi,从头预训练,性能媲美自回归模型。
Sumi: Open Uniform Diffusion Language Model from Scratch

- 首次从零开始训练大规模统一扩散语言模型,支持任意词元任意步更新。
- 在1.5万亿词上训练,知识、推理和编码任务表现接近自回归模型。
- 开源权重与完整训练方案,推动统一扩散模型研究发展。
扩散模型已成为自回归模型的有力替代。其中,统一扩散语言模型(UDLM)理论上允许任意词元在任意步骤被更新,具备更强生成灵活性。然而,目前尚无在大参数量和大词元预算下从头预训练的UDLM。自回归与掩码扩散模型已有可研究的规模模型,但统一扩散模型仍缺参考基准。为此,我们提出Sumi(日语意为“墨”),一个完全开放的70亿参数统一扩散语言模型,在1.5万亿词上从头预训练。Sumi在知识、推理和编码基准上表现与同规模自回归模型相当,但在常识任务上表现较弱,可能源于其教育类数据占比过高。我们发布模型权重、检查点及完整训练配方,包括基于公开语料的数据混合规范。希望该发布能促进社区对大规模原生统一扩散模型的研究,并激发对其尚未充分理解特性的探索。
原文摘要 · Abstract (English)
Diffusion models have become a promising alternative to autoregressive models. Among these, uniform diffusion language models (UDLMs) permit any token to be updated at any step, in principle enabling more flexible generation. However, no UDLM has yet been pretrained from scratch at both large parameter scale and large token budget. Both autoregressive modeling and masked diffusion modeling already have capable models at scale that the community can study and build on; uniform diffusion has none. A scratch-pretrained UDLM at scale would provide a clean reference point for studying scaling behavior, generation dynamics, controllability, and trade-offs against established autoregressive and masked diffusion models. To this end, we introduce Sumi ("ink" in Japanese), a fully open 7B uniform diffusion language model pretrained from scratch on 1.5T tokens. Sumi performs competitively with autoregressive models trained at comparable token budgets on knowledge, reasoning, and coding benchmarks, while under-performing on commonsense benchmarks, where our education-heavy data mixture is a likely contributor. We release our model weights, checkpoints, and full training recipe, including a complete specification of the data mixture over publicly available corpora. We hope this release enables the community to study native uniform diffusion at scale and catalyzes work on its as-yet poorly understood aspects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。