融合物理知识与人类反馈,提升自动驾驶安全性和可靠性。
Trustworthy Human-AI Collaboration: Reinforcement Learning with Human Feedback and Physics Knowledge for Safe Autonomous Driving
- 引入物理模型与人类干预协同优化决策,动态选择动作。
- 即使人类反馈质量下降,仍保证性能不低于基础物理策略。
- 适合对安全性要求高的智能驾驶系统研发者参考。
在自动驾驶领域,构建安全可信的自主驾驶策略仍是重大挑战。近期,基于人类反馈的强化学习(RLHF)因其提升训练安全性和采样效率的潜力受到广泛关注。然而,现有方法在面对不完美人类示范时易出现训练震荡,甚至性能劣于规则驱动方法。受人类学习过程启发,本文提出物理增强型人类-智能体反馈强化学习(PE-RLHF)。该框架将人类反馈(如人工干预与示范)与物理知识(如交通流模型)联合融入强化学习训练循环。核心优势在于:即使人类反馈质量下降,所学策略性能也至少不低于给定的物理基线策略,从而确保可信赖的安全性提升。PE-RLHF设计了物理增强的人机协同机制,采用无奖励代理价值函数捕捉人类偏好,并引入最小干预机制降低人类导师的认知负担。在多种驾驶场景下的大量实验表明,该方法显著优于传统方法,在安全性、效率和泛化能力上达到当前最优水平,且对不同质量的人类反馈均具鲁棒性。该思想不仅推动自动驾驶技术发展,也可为其他高安全要求领域提供启示。演示视频与代码已公开: https://zilin-huang.github.io/PE-RLHF-website/
原文摘要 · Abstract (English)
In the field of autonomous driving, developing safe and trustworthy autonomous driving policies remains a significant challenge. Recently, Reinforcement Learning with Human Feedback (RLHF) has attracted substantial attention due to its potential to enhance training safety and sampling efficiency. Nevertheless, existing RLHF-enabled methods often falter when faced with imperfect human demonstrations, potentially leading to training oscillations or even worse performance than rule-based approaches. Inspired by the human learning process, we propose Physics-enhanced Reinforcement Learning with Human Feedback (PE-RLHF). This novel framework synergistically integrates human feedback (e.g., human intervention and demonstration) and physics knowledge (e.g., traffic flow model) into the training loop of reinforcement learning. The key advantage of PE-RLHF is its guarantee that the learned policy will perform at least as well as the given physics-based policy, even when human feedback quality deteriorates, thus ensuring trustworthy safety improvements. PE-RLHF introduces a Physics-enhanced Human-AI (PE-HAI) collaborative paradigm for dynamic action selection between human and physics-based actions, employs a reward-free approach with a proxy value function to capture human preferences, and incorporates a minimal intervention mechanism to reduce the cognitive load on human mentors. Extensive experiments across diverse driving scenarios demonstrate that PE-RLHF significantly outperforms traditional methods, achieving state-of-the-art (SOTA) performance in safety, efficiency, and generalizability, even with varying quality of human feedback. The philosophy behind PE-RLHF not only advances autonomous driving technology but can also offer valuable insights for other safety-critical domains. Demo video and code are available at: \https://zilin-huang.github.io/PE-RLHF-website/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。