arXiv:2511.03492cs.LGstat.ML2025-11被引 9

提出数据精简的理论框架,解释为何少而精的数据有时比海量数据更有效。

Why Less is More (Sometimes): A Theory of Data Curation

  • 基于难易度与正确性筛选数据,构建可解析的泛化误差模型。
  • 在特定条件下,小规模精选数据集性能超越全量数据集。
  • 适用于大模型训练优化与数学推理任务中的数据策略设计。

本文提出一个理论框架,解决现代机器学习中的核心悖论:何时使用更少数据更好?随着经典缩放定律(‘更多更好’)受到挑战,如LIMO和s1等方法通过激进的数据精简实现更优性能。我们研究了由不完美预言机根据数据难度和正确性选择训练样本的策略,推导出标签无关与标签相关的测试误差缩放律,揭示了在何种条件下保留子集数据能提升泛化能力。与传统缩放律不同,我们证明在某些条件下,小规模精选数据集可超越完整数据集,并给出精确的相变曲线,关联数据规模与质量。我们在ImageNet上验证了这些理论预测,确认数据精简可提升准确率并缓解模型坍塌。此外,该框架为大型语言模型数学推理中观察到的矛盾数据策略提供了原则性解释。

原文摘要 · Abstract (English)

This paper introduces a theoretical framework to resolve a central paradox in modern machine learning: When is it better to use less data? This question has become critical as classical scaling laws suggesting ``more is more'' (Sun et al., 2025) are challenged by methods like LIMO (``less is more'') and s1 (Ye et al., 2025; Muenighoff et al., 2025), which achieve superior performance with small, aggressively curated datasets. Here, we study data curation strategies where an imperfect oracle selects the training examples according to their difficulty and correctness. Our results provide exact scaling law curves for test error under both label-agnostic and label-aware curation rules, revealing when and why keeping only a subset of data can improve generalization. In contrast to classical scaling laws, we show that under certain conditions, small curated datasets can outperform full datasets, and we provide analytical conditions for this by deriving precise phase transition curves tied to data size and quality. We validate these theoretical claims with empirical results on ImageNet, confirming our predictions about when curation improves accuracy and can even mitigate model collapse. Furthermore, our framework provides a principled explanation for the contradictory curation strategies recently observed in LLM mathematical reasoning.

数据精简泛化能力理论分析大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。