arXiv:2508.06412cs.LGcs.CL2025-08被引 4

用周期重置机制提升大模型优化样本效率,缓解过拟合问题。

Sample-efficient LLM Optimization with Reset Replay

  • 周期性重置初始数据与策略,维持网络可塑性。
  • 在数学与通用推理任务上显著提升性能,接近复杂基线效果。
  • 适合作为现有训练流程的轻量级插件,适用于资源受限场景。

近期基于强化学习和偏好优化的大模型后训练进展显著提升了其推理能力。然而,这些方法常面临样本效率低和首因偏差(primacy bias)问题,即过度拟合初始经验导致网络可塑性下降。为此,我们提出大模型重置回放优化(LoRR),一种通用且高效的偏好优化增强插件。其核心机制通过高频率回放最大化每批数据价值,并采用周期性重置策略重新利用初始数据与策略以抑制过拟合,同时结合混合优化目标更充分地利用训练数据。大量实验表明,LoRR显著提升多种偏好优化方法在数学与通用推理基准上的表现。值得注意的是,结合迭代DPO框架的LoRR在挑战性数学任务上达到与许多复杂或高计算成本基线相当的性能。结果表明,LoRR提供了一种从有限离线数据中实现高效优化的实用范式,仅需极少改动即可提升现有后训练流程性能。

原文摘要 · Abstract (English)

Recent advancements in LLM post-training, particularly through reinforcement learning and preference optimization, are key to boosting their reasoning capabilities. However, these methods often suffer from low sample efficiency and a susceptibility to primacy bias, a phenomenon where overfitting to initial experiences diminishes network plasticity and damages the learning process. To address these challenges, we introduce LLM optimization with Reset Replay (LoRR), a general and powerful plugin for enhancing sample efficiency in preference-based optimization. Its core mechanism enables high-replay training to maximize the utility of each data batch. To mitigate overfitting, LoRR orchestrates a periodic reset strategy that reuses the initial data and policy to maintain network plasticity, and further adopts a hybrid optimization objective to better exploit training data. Extensive experiments show that LoRR significantly boosts the performance of various preference optimization methods on both mathematical and general reasoning benchmarks. Notably, an iterative DPO framework augmented with LoRR achieves comparable performance on challenging math tasks, rivaling many complex or computationally expensive baselines. Our findings highlight that LoRR offers a practical and sample-efficient paradigm from limited offline data, unlocking greater performance with minimal changes to existing post-training workflows.

大模型优化样本效率偏好学习重置机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。