对比两种模型在小数据下的效率,发现自回归模型更省数据。
Double Descent as a Lens for Sample Efficiency in Autoregressive vs. Discrete Diffusion Models
- 用双下降现象分析模型样本效率,比较自回归与离散扩散模型。
- 小数据下自回归模型表现更好,扩散模型需更大容量和更多训练才可追上。
- 适合关注模型高效训练、数据有限场景的研究者阅读。
数据稀缺推动对更高效大语言模型的需求。本文利用双下降现象,全面比较离散扩散模型与自回归模型的样本效率。结果表明,离散扩散模型需要更大的模型容量和更多的训练轮次才能脱离欠参数化阶段并达到插值阈值。在强过参数化阶段,两类模型表现相似,且在大范围模型规模下均未出现明显的第二次测试损失下降。总体而言,自回归模型在小规模数据集上更具样本效率;而离散扩散模型仅在具备充足容量和计算资源时才具有竞争力。
原文摘要 · Abstract (English)
Data scarcity drives the need for more sample-efficient large language models. In this work, we use the double descent phenomenon to holistically compare the sample efficiency of discrete diffusion and autoregressive models. We show that discrete diffusion models require larger capacity and more training epochs to escape their underparameterized regime and reach the interpolation threshold. In the strongly overparameterized regime, both models exhibit similar behavior, with neither exhibiting a pronounced second descent in test loss across a large range of model sizes. Overall, our results indicate that autoregressive models are more sample-efficient on small-scale datasets, while discrete diffusion models only become competitive when given sufficient capacity and compute.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。