奖励设计直接影响自动驾驶智能体关注什么,决定其安全感知策略。
Reward-Conditioned Attention: How Reward Design Shapes What Autonomous Driving Agents See

- 通过对比不同奖励机制,发现奖励类型决定智能体对道路要素的关注重点。
- 导航奖励使智能体对路径标记的关注度最高达4.7倍于无导航激励模型。
- 连续碰撞时间惩罚带来持续警觉状态,提升全程监控能力。
我们研究了奖励设计如何影响强化学习自动驾驶智能体的内部注意力模式。采用三个基于Perceiver的智能体,它们拥有相同架构与训练数据,仅奖励配置不同——从基础违规惩罚到连续接近惩罚。在Waymo Open Motion Dataset中的50个真实场景下分析跨注意力分配。方法论发现:简单的时间步聚合会低估注意力与风险的关系;使用组内事件的Fisher z变换聚合才是恰当统计,揭示出碰撞风险与注意力呈稳健正相关。基于此方法,发现:带导航奖励的智能体对GPS路径标记的关注度最高可达无导航激励模型的4.7倍,以及带邻近惩罚模型的2.0倍;连续时间到碰撞惩罚催生“学习型警觉先验”,即在无碰撞阶段也维持高监控水平。部分场景中,完整奖励与最小奖励模型的注意力-风险相关方向相反,表明奖励设计可彻底反转注意力策略,而不仅是调节强度。结果表明,注意力分析是验证奖励函数是否实现预期表征行为的实用诊断工具。
原文摘要 · Abstract (English)
We investigate how reward design shapes the internal attention patterns of reinforcement learning agents trained for autonomous driving. Using three Perceiver-based agents that share identical architectures and training data but differ only in their reward configurations$\unicode{x2014}$ranging from basic violation penalties to continuous proximity penalties$\unicode{x2014}$we analyze cross-attention allocation across 50 real-world scenarios from the Waymo Open Motion Dataset. A central methodological finding is that naïve pooling of timesteps across episodes substantially underestimates the attention$\unicode{x2013}$risk relationship; within-episode correlation with Fisher z-transform aggregation is the appropriate statistic and reveals a robustly positive link between collision risk and agent-directed attention. Building on this validated methodology, we demonstrate two reward-conditioned effects: agents trained with navigation rewards allocate up to $2.0\times$ more attention to GPS-path tokens than those trained with additional proximity penalties$\unicode{x2014}$and $4.7\times$ more than agents with no navigation incentive$\unicode{x2014}$revealing that reward content directly determines which scene elements the encoder prioritizes, and continuous time-to-collision penalties create a $\textit{learned vigilance prior}$$\unicode{x2014}$elevated resting agent surveillance maintained throughout collision-free phases. In several scenarios, the complete-reward and minimal-reward models exhibit opposite attention$\unicode{x2013}$risk correlation directions, demonstrating that reward design can qualitatively reverse attentional strategy rather than merely modulating its magnitude. These results suggest that attention analysis is a practical diagnostic for verifying that a reward function produces the intended representational behaviour in safety-critical RL systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。