用频率感知掩码让语言模型在少数据下高效学习。
Masked Diffusion Language Models with Frequency-Informed Training
- 基于扩散模型,按词频动态掩码,优先学罕见词。
- 在BabyLM基准上表现媲美混合自回归模型。
- 适合数据稀缺场景下的语言模型训练。
我们提出一种针对BabyLM 2025挑战的掩码扩散语言建模框架,旨在实现严格数据约束下的高效训练。该方法将扩散训练目标应用于语言建模,引入频率感知掩码策略,优先学习稀有词汇,同时保持理论有效性。我们探索了多种噪声调度策略,包括双模式方法,并研究了NELBO目标内不同的噪声加权方案。在BabyLM基准套件上评估了该方法,衡量语言能力、世界知识和人类化程度。结果表明,性能可与混合自回归-掩码基线相媲美,证明扩散训练在数据受限条件下具有可行性。
原文摘要 · Abstract (English)
We present a masked diffusion language modeling framework for data-efficient training for the BabyLM 2025 Challenge. Our approach applies diffusion training objectives to language modeling under strict data constraints, incorporating frequency-informed masking that prioritizes learning from rare tokens while maintaining theoretical validity. We explore multiple noise scheduling strategies, including two-mode approaches, and investigate different noise weighting schemes within the NELBO objective. We evaluate our method on the BabyLM benchmark suite, measuring linguistic competence, world knowledge, and human-likeness. Results show performance competitive to hybrid autoregressive-masked baselines, demonstrating that diffusion-based training offers a viable alternative for data-restricted language learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。