让智能体学会自我怀疑与恢复,提升强化学习在噪声环境下的稳定性。
Meta-Cognitive Reinforcement Learning with Self-Doubt and Recovery
- 引入内部可信度信号,动态调节学习行为
- 在奖励污染环境下平均收益更高,后期训练失败率显著降低
- 适合需要高可靠性的机器人控制等真实场景
稳健的强化学习方法通常关注抑制不可靠经验或受损奖励,但缺乏对自身学习过程可靠性的推理能力。因此,这些方法要么对噪声过度反应而变得过于保守,要么在不确定性累积时发生灾难性失败。本文提出一种元认知强化学习框架,使智能体能够基于内部估计的可靠性信号评估、调控并恢复其学习行为。该方法引入由价值预测误差稳定性(VPES)驱动的元信任变量,通过故障保护机制和渐进式信任恢复来调节学习动态。在带奖励污染的连续控制基准测试中,具备恢复能力的元认知控制相比强基线方法实现了更高的平均回报,并显著减少了训练后期的失败次数。
原文摘要 · Abstract (English)
Robust reinforcement learning methods typically focus on suppressing unreliable experiences or corrupted rewards, but they lack the ability to reason about the reliability of their own learning process. As a result, such methods often either overreact to noise by becoming overly conservative or fail catastrophically when uncertainty accumulates. In this work, we propose a meta-cognitive reinforcement learning framework that enables an agent to assess, regulate, and recover its learning behavior based on internally estimated reliability signals. The proposed method introduces a meta-trust variable driven by Value Prediction Error Stability (VPES), which modulates learning dynamics via fail-safe regulation and gradual trust recovery. Experiments on continuous-control benchmarks with reward corruption demonstrate that recovery-enabled meta-cognitive control achieves higher average returns and significantly reduces late-stage training failures compared to strong robustness baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。