arXiv:2505.24317cs.LG2025-05

让自动驾驶模型懂交通规则,减少事故责任。

ROAD: Responsibility-Oriented Reward Design for Reinforcement Learning in Autonomous Driving

  • 用知识图谱+大模型自动设计奖励函数
  • 事故责任判定准确率显著提升
  • 适合追求安全合规的自动驾驶研发

自动驾驶中的强化学习依赖试错机制,在复杂环境中提升鲁棒性。然而,传统奖励函数多依赖人工设计,难以应对复杂场景。本文提出一种责任导向的奖励设计方法,将交通法规显式融入强化学习框架。通过构建交通法规知识图谱,并结合视觉-语言模型与检索增强生成技术,实现奖励的自动化分配。该方法引导智能体严格遵守交通规则,有效降低违规行为,优化决策性能。实验表明,该方法显著提升了事故责任判定的准确性,同时大幅减少了智能体在交通事故中的责任风险。

原文摘要 · Abstract (English)

Reinforcement learning (RL) in autonomous driving employs a trial-and-error mechanism, enhancing robustness in unpredictable environments. However, crafting effective reward functions remains challenging, as conventional approaches rely heavily on manual design and demonstrate limited efficacy in complex scenarios. To address this issue, this study introduces a responsibility-oriented reward function that explicitly incorporates traffic regulations into the RL framework. Specifically, we introduced a Traffic Regulation Knowledge Graph and leveraged Vision-Language Models alongside Retrieval-Augmented Generation techniques to automate reward assignment. This integration guides agents to adhere strictly to traffic laws, thus minimizing rule violations and optimizing decision-making performance in diverse driving conditions. Experimental validations demonstrate that the proposed methodology significantly improves the accuracy of assigning accident responsibilities and effectively reduces the agent's liability in traffic incidents.

强化学习自动驾驶奖励设计交通规则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。