动态调整奖励权重与数据重要性,提升大模型对复杂任务的适应能力。
SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility
- 根据学习进展自动调节多目标奖励权重和数据优先级。
- 在多个基准上显著提升模型在各领域的能力表现。
- 适合需要持续优化对齐策略的复杂真实场景应用。
大型语言模型的发展正从单一可验证任务转向复杂的开放世界场景,给后训练阶段带来巨大挑战。在此类场景中,奖励系统的规模与复杂性显著增加,趋向于涵盖模型多种能力与应用场景的多目标形式。然而,传统方法通常采用固定奖励权重,忽视了非平稳的学习动态,并难以应对跨维度的数据异质性。为此,我们提出SPARD框架,通过感知学习进度,自动建立自适应课程,动态调整多目标奖励权重与数据重要性,从而实现学习意图与数据效用的同步优化,以达到最佳性能。在多个基准上的大量实验表明,SPARD显著提升了模型在所有领域的综合能力。
原文摘要 · Abstract (English)
The evolution of Large Language Models (LLMs) is shifting the focus from single, verifiable tasks toward complex, open-ended real-world scenarios, imposing significant challenges on the post-training phase. In these settings, the scale and complexity of reward systems have grown significantly, transitioning toward multi-objective formulations that encompass a comprehensive spectrum of model capabilities and application contexts. However, traditional methods typically rely on fixed reward weights, ignoring non-stationary learning dynamics and struggling with data heterogeneity across dimensions. To address these issues, we propose SPARD, a framework that establishes an automated, self-paced curriculum by perceiving learning progress to dynamically adjust multi-objective reward weights and data importance, thereby synchronizing learning intent with data utility for optimal performance. Extensive experiments across multiple benchmarks demonstrate that SPARD significantly enhances model capabilities across all domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。