arXiv:2510.06038cs.LGcs.AI2025-10被引 3

人类引导强化学习让自动驾驶更安全高效

From Learning to Mastery: Achieving Safe and Efficient Real-World Autonomous Driving with Human-In-The-Loop Reinforcement Learning

  • 用人类反馈构建代理价值函数,指导智能体学习专家行为
  • 在真实驾驶环境中实现高效、安全的策略训练,无需预先定义奖励函数
  • 适合需要高安全性与样本效率的自动驾驶系统研发者

基于强化学习的自动驾驶具有巨大潜力,但在真实场景中应用仍面临安全、高效和鲁棒性挑战。通过引入人类专家知识,可减少危险探索并提升样本效率。本文提出一种无奖励、主动式人机协同学习方法——人类引导分布软演员-批评家(H-DSAC)。该方法结合代理价值传播(PVP)与分布软演员-批评家(DSAC),在DSAC框架内构建分布式代理价值函数,通过赋予专家示范更高预期回报并惩罚需人工干预的动作,编码人类意图。该函数将标签外推至未标注状态,有效引导策略向专家行为靠拢。在精心设计的状态空间下,方法可在实际训练时间内完成真实世界驾驶策略学习。仿真与真实实验结果均表明,该框架实现了自动驾驶的安全、鲁棒且高效的样本学习。

原文摘要 · Abstract (English)

Autonomous driving with reinforcement learning (RL) has significant potential. However, applying RL in real-world settings remains challenging due to the need for safe, efficient, and robust learning. Incorporating human expertise into the learning process can help overcome these challenges by reducing risky exploration and improving sample efficiency. In this work, we propose a reward-free, active human-in-the-loop learning method called Human-Guided Distributional Soft Actor-Critic (H-DSAC). Our method combines Proxy Value Propagation (PVP) and Distributional Soft Actor-Critic (DSAC) to enable efficient and safe training in real-world environments. The key innovation is the construction of a distributed proxy value function within the DSAC framework. This function encodes human intent by assigning higher expected returns to expert demonstrations and penalizing actions that require human intervention. By extrapolating these labels to unlabeled states, the policy is effectively guided toward expert-like behavior. With a well-designed state space, our method achieves real-world driving policy learning within practical training times. Results from both simulation and real-world experiments demonstrate that our framework enables safe, robust, and sample-efficient learning for autonomous driving.

自动驾驶强化学习人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。