幂律分布比均匀分布更利于模型学习长尾组合技能。
The Power of Power Law: Asymmetry Enables Compositional Reasoning

- 用幂律分布采样数据,显著降低训练所需样本量。
- 在多步算术等任务中,幂律训练比均匀分布性能更优。
- 适合研究数据分布与模型泛化能力的学者。
自然语言数据服从幂律分布,多数知识与技能出现频率极低。尽管普遍认为重新加权或整理数据至均匀分布有助于模型学习长尾技能,但我们发现:在多种组合推理任务(如状态追踪、多步算术)中,幂律分布训练始终优于均匀分布。通过设计一个极简的技能组合任务,我们证明幂律采样可显著减少训练数据需求。理论分析表明,幂律采样引入有益的不对称性,改善了病态损失曲面,使模型先以低数据复杂度掌握高频技能组合,进而高效学习稀有长尾技能。研究为有效数据分布提供了新视角。
原文摘要 · Abstract (English)
Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intuition suggests that reweighting or curating data towards a uniform distribution may help models better learn these long-tail skills, we find a counterintuitive result: across a wide range of compositional reasoning tasks, such as state tracking and multi-step arithmetic, training under power-law distributions consistently outperforms training under uniform distributions. To understand this advantage, we introduce a minimalist skill-composition task and show that learning under a power-law distribution provably requires significantly less training data. Our theoretical analysis reveals that power law sampling induces a beneficial asymmetry that improves the pathological loss landscape, which enables models to first acquire high-frequency skill compositions with low data complexity, which in turn serves as a stepping stone to efficiently learn rare long-tailed skills. Our results offer an alternative perspective on what constitutes an effective data distribution for training models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。