arXiv:2511.20591cs.LG2025-11

通过注意力轨迹分析,揭示强化学习模型隐藏的决策偏好与漏洞。

Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning

  • 用分层注意力图谱追踪模型对输入特征的关注变化过程。
  • 发现不同算法有特定注意力偏倚,且与异常策略和过拟合相关。
  • 适合研究模型可解释性、训练缺陷诊断的AI从业者。

尽管深度强化学习代理在多个领域表现优异,但仅通过性能指标难以理解其内部决策过程。我们提出一种科学方法,通过量化显著性分析学习过程。该方法将对象级和模态级显著性信息聚合为层次化注意力分布,量化代理随时间分配注意力的方式,形成训练过程中的注意力轨迹。应用于Atari基准、自定义Pong环境及视觉运动交互任务中的肌动生物力学用户仿真,该方法揭示了算法特有的注意力偏倚,暴露了由奖励驱动的非预期策略,并诊断出对冗余感官通道的过拟合。这些模式与可测量的行为差异对应,证明注意力分布、学习动态与代理行为之间存在实证关联。通过多种显著性方法和环境验证了注意力分布的鲁棒性。结果表明,注意力轨迹是追踪特征依赖演化、识别性能指标无法察觉的偏倚与脆弱性的有力诊断工具。

原文摘要 · Abstract (English)

While deep reinforcement learning agents demonstrate high performance across domains, their internal decision processes remain difficult to interpret when evaluated only through performance metrics. In particular, it is poorly understood which input features agents rely on, how these dependencies evolve during training, and how they relate to behavior. We introduce a scientific methodology for analyzing the learning process through quantitative analysis of saliency. This approach aggregates saliency information at the object and modality level into hierarchical attention profiles, quantifying how agents allocate attention over time, thereby forming attention trajectories throughout training. Applied to Atari benchmarks, custom Pong environments, and muscle-actuated biomechanical user simulations in visuomotor interactive tasks, this methodology uncovers algorithm-specific attention biases, reveals unintended reward-driven strategies, and diagnoses overfitting to redundant sensory channels. These patterns correspond to measurable behavioral differences, demonstrating empirical links between attention profiles, learning dynamics, and agent behavior. To assess robustness of the attention profiles, we validate our findings across multiple saliency methods and environments. The results establish attention trajectories as a promising diagnostic axis for tracing how feature reliance develops during training and for identifying biases and vulnerabilities invisible to performance metrics alone.

可解释性强化学习注意力分析模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。