arXiv:2605.20314cs.LGcs.AI2026-05

小数据重复训练更快,因采样偏差促进分层优化。

Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases

论文配图:Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases
图 1 · 摘自论文原文
  • 通过重复小数据集,利用采样偏差实现分层优化
  • 小数据重复可节省计算量,加速训练过程
  • 特别适合推理类任务的主动优化策略

本文研究了「小数据与大数据之间的差距」现象:在某些算法任务、模型架构和优化器下,重复少量样本比使用更大数据集更能节省计算资源。该现象无法用现有理论解释。我们提出,加速源于采样偏差带来的合适分层增长,且在数据量较小时更为显著。通过理论分析和多种干预实验,验证了这一机制。结果表明,在数据稀缺时重复小数据并非权宜之计,而是一种可主动利用的有利归纳偏置,尤其适用于推理类任务。

原文摘要 · Abstract (English)

This work investigates the ``small-vs-large gap'', where repeating on fewer samples can lead to compute saving during training compared to using a larger dataset. This is observed across algorithmic tasks, architectures and optimizers and cannot be explained using prior theory. We argue that the speedup comes from appropriate layer-wise growth enabled by sampling biases, which is more pronounced when the dataset size is smaller. We provide both theoretical analysis and empirical evidence from various interventions. Our results suggest that using a smaller dataset with more repetitions is not just a fallback strategy under data scarcity, but can be proactively leveraged as a favorable inductive biases for optimization, particularly in reasoning tasks.

训练加速采样偏差小数据优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。