arXiv:2607.04969cs.LGcs.CL2026-07

通过记忆窗口信号优化数据重用,让大模型训练更高效

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training

  • 基于损失保留与评估分数动态识别模型记忆窗口
  • 实验显示重复训练远超现有四轮标准仍有效提升性能
  • 适合追求样本效率的高效大模型训练团队参考

大语言模型训练正从单遍训练转向多轮训练,合理重用有限高质量数据可提升模型表现与样本效率。但过度重复会引发过拟合与收益递减。本文通过观察模型‘记忆窗口’信号(来自损失保留动态与下游评估分数),提出‘记忆引导的数据重用’训练范式,自适应决定数据重用时机与方式,实现对训练轮数和数据重播调度的合理决策。初步实验揭示一致的记忆驱动规律:性能在远超当前实践(如常见的四轮限制)后仍持续提升。尽管完整调度器尚待未来工作,这些发现为记忆感知训练调度提供了基础,有助于确定数据重用预算,推动在有限高质量数据下‘更聪明地训练’而非‘更长时间训练’。

原文摘要 · Abstract (English)

The training paradigm of large language models has shifted from traditional one-pass training to multi-epoch training, as reasonable reuse of limited high-quality data can improve both model performance and sample efficiency. Meanwhile, excessive repetition introduces the risk of overfitting and diminishing returns. Determining when and how to reuse data effectively thus emerges as a natural but under-explored question. Through a novel observation of model's "Memorization Window" signals derived from loss retention dynamics and downstream evaluation scores, we propose "Memorization-guided Data Reuse", a training paradigm that adaptively determines when and how data should be reused, enabling principled decisions on the number of training epochs and the scheduling of data replays. Our preliminary experiments reveal a consistent memorization-driven regime: performance continues to improve with repetition far beyond current practice (e.g., the commonly cited four-epoch limit). While a full scheduler remains future work, these insights provide a foundation for memorization-aware training schedules, helping to determine reuse budgets and move toward training LLMs smarter rather than longer with limited high-quality data.

大模型训练数据重用样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。