用策略感知惊喜提升强化学习探索效率
SuS: Strategy-aware Surprise for Intrinsic Exploration
- 基于策略稳定性与策略意外性双信号设计新探索机制
- 在数学推理任务上提升17.4%的Pass@1和26.4%的Pass@5
- 适合需要高多样性解法的复杂决策场景
我们提出策略感知惊喜(SuS),一种新颖的内在动机框架,利用预-后预测差异作为强化学习中的探索新颖性信号。不同于仅依赖状态预测误差的传统好奇心方法,SuS引入两个互补组件:策略稳定性(SS)衡量行为策略在时间步上的一致性,策略惊喜(SuS)捕捉相对于当前策略表征的意外结果。联合奖励通过学习的权重系数融合两者信号。我们在大语言模型的数学推理任务上评估了SuS,显著提升了准确率与解法多样性。消融实验表明,移除任一组件均导致至少10%的性能下降,验证了方法的协同效应。相比基线,SuS在Pass@1上提升17.4%,在Pass@5上提升26.4%,且训练全程保持更高策略多样性。
原文摘要 · Abstract (English)
We propose Strategy-aware Surprise (SuS), a novel intrinsic motivation framework that uses pre-post prediction mismatch as a novelty signal for exploration in reinforcement learning. Unlike traditional curiosity-driven methods that rely solely on state prediction error, SuS introduces two complementary components: Strategy Stability (SS) and Strategy Surprise (SuS). SS measures consistency in behavioral strategy across temporal steps, while SuS captures unexpected outcomes relative to the agent's current strategy representation. Our combined reward formulation leverages both signals through learned weighting coefficients. We evaluate SuS on mathematical reasoning tasks using large language models, demonstrating significant improvements in both accuracy and solution diversity. Ablation studies confirm that removing either component results in at least 10% performance degradation, validating the synergistic nature of our approach. SuS achieves 17.4% improvement in Pass@1 and 26.4% improvement in Pass@5 compared to baseline methods, while maintaining higher strategy diversity throughout training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。