arXiv:2605.07820cs.LG2026-05被引 1

17亿参数流模型4步生成高质量文本,突破离散数据连续建模规模瓶颈。

Scaling Categorical Flow Maps

论文配图:Scaling Categorical Flow Maps
图 1 · 摘自论文原文
  • 用流匹配在高维离散空间构建连续路径,实现快速采样。
  • 1.7B模型在2.1万亿词上训练,4步生成文本质量媲美长序列采样。
  • 提出半离散似然界,可评估模型性能,适合大规模语言建模研究者。

连续扩散与流匹配模型为语言建模提供了有潜力的替代方案,使离散数据能享受连续建模的优势,如加速采样和分布倾斜。近期工作已证明通过在高斯分布与独热编码数据分布间进行简单流匹配,可实现离散数据的连续生成,并利用分类流映射(CFM)实现加速采样,在少步数下达到良好样本质量。然而此前仅在参数量小于10亿的模型上验证,其可扩展性仍未知。本文在2.1万亿词数据上训练了17亿参数的基础流模型,并自蒸馏为CFM,可在4步内生成多样化、高质量文本,同时保持接近数据级的词元熵。我们还提出了半离散设置下的似然界,可用于标准语言建模基准评分,结果与离散扩散方法相当。最后,我们揭示了大规模训练中的挑战,给出了损失加权与时间调度的实用建议。

原文摘要 · Abstract (English)

Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they unlock a host of advantages currently reserved for continuous modalities, including accelerated sampling and tilting. Recently, several works have demonstrated the possibility of generating discrete data continuously by a simple flow matching process between a Gaussian and the one-hot encoded data distribution. They have further shown the feasibility of accelerated sampling via Categorical Flow Maps (CFMs), resulting in competitive sample quality in the few-step regime. However, this method had only been evaluated at relatively modest scales ($<1$B), leaving the question of its scalability completely open. In this article, we train a $1.7$B-parameter base flow model on $2.1$T tokens and self-distill it into a CFM that generates diverse, high-quality text in as few as $4$ inference steps while maintaining near-data-level token entropy. Furthermore, we introduce a likelihood bound for CFMs in the semi-discrete setting, and show that they can be used to score the model on standard LM benchmarks, achieving results in the same range as discrete diffusion methods. Finally, we uncover some of the challenges that arise from training these models at scale, and we provide prescriptive insights on loss weighting and time scheduling.

流模型语言建模高效生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。