通过持续一致性机制提升语言模型强化学习的训练稳定性。
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

- 基于教师信号的局部持续性动态生成蒸馏权重
- 在ALFWorld上超越GRPO 15.6和13.3分,未使用推理技能
- 适合追求高效强化学习训练的语言模型研究者
大型语言模型代理在复杂交互任务中展现强大潜力,但其强化学习常受稀疏奖励制约,长多轮轨迹仅获单一结果信号。现有策略依赖孤立的词元级差异或共享步级权重,易受噪声干扰或忽略位置变化。本文提出持久一致性自蒸馏(PCSD),从教师偏好信号的局部持续性中推导词元级蒸馏权重。通过自适应窗口与指数衰减聚合捕捉持续的教师支持,结合趋势感知调制抑制局部下降支持,并用Sigmoid门控生成连续权重。该目标与GRPO联合优化,融合密集教师指导与稀疏环境反馈。无需推理时技能,PCSD在两个骨干模型上均取得最佳ALFWorld综合表现,分别优于GRPO 15.6和13.3分,优于SDAR 6.2和5.5分,在WebShop上保持竞争力,并在未见ALFWorld划分上比GRPO提升15.8分。
原文摘要 · Abstract (English)
Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may receive only a single outcome-level signal. On-policy self-distillation (OPSD) provides dense token-level supervision from a privileged teacher, but the teacher may not be reliable at every position. Existing methods commonly rely on isolated token-level discrepancies, which can be sensitive to noise, or assign a shared step-level weight that may overlook positional variation. We propose Persistent Consistency Self-Distillation (PCSD), which derives token-level distillation weights from the local persistence of teacher-favoring signals. PCSD combines adaptive windows with exponentially decayed aggregation to capture persistent relative teacher support, applies trend-aware modulation to attenuate locally declining support, and produces continuous weights through sigmoid gating. The resulting objective is jointly optimized with GRPO, combining dense teacher guidance with sparse environmental feedback. Without inference-time skills, PCSD achieves the best ALFWorld Overall results among all baselines on both backbones, exceeding GRPO by 15.6 and 13.3 points and SDAR by 6.2 and 5.5 points, while remaining competitive on WebShop and gaining 15.8 points over GRPO on unseen ALFWorld split.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。