arXiv:2606.04735cs.LGcs.AI2026-06

发现深度强化学习中奖励峰值偏误,解释人类记忆偏差的根源

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

  • 通过归因痕迹机制揭示奖励峰值偏好现象
  • 中间迹深时高尖峰奖励被优先选择,即使总回报更低
  • 自适应优化器可缓解此偏差,适合研究认知与智能的学者

时间信用分配是生物与人工智能的核心问题,但其与非线性函数逼近的交互仍不清晰。我们识别出深度强化学习中一种系统性失效模式——迹媒介峰值偏误(TMPB)。在中等归因痕迹深度下,智能体不合理地偏好具有高幅度奖励‘峰值’的轨迹,而非累积回报更高的路径。这为‘峰终定律’(人类记忆中以最强烈时刻评价整体体验)提供了机制解释。我们发现,归因痕迹会将远端时间差分误差放大为‘梯度冲击’,固定步长随机梯度下降无法调节,导致全局值估计过高。相反,自适应优化器通过二阶矩归一化可缓解该病理。结果表明,人类式的显著性扭曲可能源于分布式系统中信用分配的数学约束,而自适应优化是实现理性价值估计的理论必要条件。

原文摘要 · Abstract (English)

Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is poorly understood. We identify a systematic failure mode in deep reinforcement learning (RL) termed Trace-Mediated Peak Bias (TMPB). At intermediate eligibility trace depths, agents irrationally prefer trajectories with high-magnitude reward ``peaks'' over alternatives with higher cumulative returns. This provides a mechanistic account of the Peak-End Rule: a human memory bias where experiences are judged by their most intense moments rather than integrated utility. We show that TMPB emerges because traces amplify distal Temporal Difference errors into ``gradient shocks'' that fixed-step-size Stochastic Gradient Descent cannot normalize, leading to global overestimation. Conversely, adaptive optimizers mitigate this pathology via second-moment normalization. Our results suggest that human-like saliency distortions may emerge naturally from the mathematical constraints of credit assignment in distributed systems, and that adaptive optimization is a theoretical necessity for rational value estimation.

强化学习信用分配认知偏差优化器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。