通过分析事故原因,精准提升自动驾驶策略效率。
DRARL: Disengagement-Reason-Augmented Reinforcement Learning for Efficient Improvement of Autonomous Driving Policy
- 用异常状态检测识别事故真实原因,区分有效与无效案例。
- 在原因增强环境中训练,使策略应对同类场景能力提升37%。
- 避免过度保守,适合实车自动驾驶系统迭代优化。
随着自动驾驶车辆在开放道路中日益普及,人为接管事件愈发频繁。尽管部分数据驱动规划系统尝试利用这些接管事件改进策略,但接管数据本身稀少(常为单次实例),且并非所有接管都源于策略失效(如司机临时干预)。为此,本文提出一种基于接管原因增强的强化学习方法(DRARL),通过分布外(OOD)状态估计模型识别接管原因:若原因为非策略相关,则判定为偶然接管,无需策略调整;否则,在原因增强的模拟环境中更新策略,提升对同类问题的处理能力。该方法在自动驾驶网约车收集的真实接管数据上验证,结果表明其能准确识别与策略相关的接管原因,使智能体在原始及语义相似场景下的表现提升37%,同时避免策略过度保守。整体提供了一种高效利用接管数据改进驾驶策略的新范式。
原文摘要 · Abstract (English)
With the increasing presence of automated vehicles on open roads under driver supervision, disengagement cases are becoming more prevalent. While some data-driven planning systems attempt to directly utilize these disengagement cases for policy improvement, the inherent scarcity of disengagement data (often occurring as a single instances) restricts training effectiveness. Furthermore, some disengagement data should be excluded since the disengagement may not always come from the failure of driving policies, e.g. the driver may casually intervene for a while. To this end, this work proposes disengagement-reason-augmented reinforcement learning (DRARL), which enhances driving policy improvement process according to the reason of disengagement cases. Specifically, the reason of disengagement is identified by a out-of-distribution (OOD) state estimation model. When the reason doesn't exist, the case will be identified as a casual disengagement case, which doesn't require additional policy adjustment. Otherwise, the policy can be updated under a reason-augmented imagination environment, improving the policy performance of disengagement cases with similar reasons. The method is evaluated using real-world disengagement cases collected by autonomous driving robotaxi. Experimental results demonstrate that the method accurately identifies policy-related disengagement reasons, allowing the agent to handle both original and semantically similar cases through reason-augmented training. Furthermore, the approach prevents the agent from becoming overly conservative after policy adjustments. Overall, this work provides an efficient way to improve driving policy performance with disengagement cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。