arXiv:2504.08704cs.ROcs.LG2025-04被引 3

用人类安全判断生成奖励标签,提升复杂路口避障能力

Offline Reinforcement Learning using Human-Aligned Reward Labeling for Autonomous Emergency Braking in Occluded Pedestrian Crossing

  • 基于语义分割动态激活安全奖励,模拟人类驾驶优先级
  • 在模拟数据上训练出有效避障策略,安全性能显著提升
  • 可发现人工标注遗漏的危险状态,适合自动驾驶安全研究

真实驾驶数据对自动驾驶系统训练至关重要。离线强化学习可利用此类数据,但多数数据缺乏有意义的奖励标签。奖励标签提供行为优劣反馈,直接影响策略性能。本文提出一种生成人类对齐奖励标签的新方法,通过分析语义分割图动态激活安全组件,在潜在碰撞场景中优先考虑安全性。该方法应用于遮挡行人过街场景,涵盖不同行人流量水平,使用仿真数据进行验证。使用生成奖励训练多个离线强化学习算法均获得有效策略,证明方法可行性。此外,在奥迪自动驾驶数据集子集上应用该方法,并与人工标注奖励对比,结果显示两组奖励存在中等差异,且该方法识别出人工标注遗漏的危险状态。

原文摘要 · Abstract (English)

Effective leveraging of real-world driving datasets is crucial for enhancing the training of autonomous driving systems. While Offline Reinforcement Learning enables training autonomous vehicles with such data, most available datasets lack meaningful reward labels. Reward labeling is essential as it provides feedback for the learning algorithm to distinguish between desirable and undesirable behaviors, thereby improving policy performance. This paper presents a novel approach for generating human-aligned reward labels. The proposed approach addresses the challenge of absent reward signals in the real-world datasets by generating labels that reflect human judgment and safety considerations. The reward function incorporates an adaptive safety component that is activated by analyzing semantic segmentation maps, enabling the autonomous vehicle to prioritize safety over efficiency in potential collision scenarios. The proposed method is applied to an occluded pedestrian crossing scenario with varying pedestrian traffic levels, using simulation data. When the generated rewards were used to train various Offline Reinforcement Learning algorithms, each model produced a meaningful policy, demonstrating the method's viability. In addition, the method was applied to a subset of the Audi Autonomous Driving Dataset, and the reward labels were compared to human-annotated reward labels. The findings show a moderate disparity between the two reward sets, and, most interestingly, the method flagged unsafe states that the human annotator missed.

离线强化学习自动驾驶安全奖励语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。