arXiv:2505.06737cs.ROcs.AI2025-05中稿 · the 36th IEEE Inte…被引 2

为自动驾驶强化学习设计更安全的奖励函数,降低碰撞率21%。

Balancing Progress and Safety: A Novel Risk-Aware Objective for RL in Autonomous Driving

  • 构建分层驱动目标并归一化权重,提升奖励透明度。
  • 引入二维椭球风险函数,显式建模碰撞前风险行为。
  • 在无信号交叉口测试中,兼顾安全与行驶效率,适合高风险场景应用。

强化学习(RL)因其强大的决策能力,被视为实现自动驾驶的有前景方法。传统方法通过试错训练驾驶策略,依赖融合多种目标的奖励函数,但该函数设计常被忽视,导致奖励定义模糊且存在缺陷。特别是安全仅以碰撞惩罚形式体现,未考虑碰撞前的行为风险,限制了实际应用。为此,本文提出分层驱动目标结构,并采用归一化方式明确各目标贡献。进一步引入基于二维椭球函数的风险感知目标,扩展责任敏感安全(RSS)概念,用于建模复杂交通交互中的风险。在不同交通密度的无信号交叉口场景中评估表明,相比基线奖励,本方法平均降低21%碰撞率,同时在路线进展和累积奖励上持续领先,证明其能有效促进安全驾驶且保持高性能。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) is a promising approach for achieving autonomous driving due to robust decision-making capabilities. RL learns a driving policy through trial and error in traffic scenarios, guided by a reward function that combines the driving objectives. The design of such reward function has received insufficient attention, yielding ill-defined rewards with various pitfalls. Safety, in particular, has long been regarded only as a penalty for collisions. This leaves the risks associated with actions leading up to a collision unaddressed, limiting the applicability of RL in real-world scenarios. To address these shortcomings, our work focuses on enhancing the reward formulation by defining a set of driving objectives and structuring them hierarchically. Furthermore, we discuss the formulation of these objectives in a normalized manner to transparently determine their contribution to the overall reward. Additionally, we introduce a novel risk-aware objective for various driving interactions based on a two-dimensional ellipsoid function and an extension of Responsibility-Sensitive Safety (RSS) concepts. We evaluate the efficacy of our proposed reward in unsignalized intersection scenarios with varying traffic densities. The approach decreases collision rates by 21\% on average compared to baseline rewards and consistently surpasses them in route progress and cumulative reward, demonstrating its capability to promote safer driving behaviors while maintaining high-performance levels.

强化学习自动驾驶安全驾驶奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。