arXiv:2509.04712cs.ROcs.AI2025-09被引 1

用非专家驾驶策略引导强化学习,提升自动驾驶训练效率。

Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving

  • 用规则式变道控制器辅助SAC算法,改善探索能力。
  • 在仿真环境中实现更稳定高效的驾驶策略学习。
  • 适合缺乏高质量示范数据的自动驾驶强化学习场景。

基于强化学习(RL)的自动驾驶控制因能通过环境交互自主学习驾驶策略而受到广泛关注。然而,RL代理常面临样本效率低和有效探索困难的问题,难以发现最优驾驶策略。为解决这些问题,我们提出使用非高度优化或非专家级的示范策略来引导RL驾驶代理。具体而言,将基于规则的变道控制器与软演员-评论家(SAC)算法结合,以增强探索能力和学习效率。实验表明,该方法显著提升了驾驶性能,且可推广至其他可受益于示范引导的驾驶场景。

原文摘要 · Abstract (English)

Automated vehicle control using reinforcement learning (RL) has attracted significant attention due to its potential to learn driving policies through environment interaction. However, RL agents often face training challenges in sample efficiency and effective exploration, making it difficult to discover an optimal driving strategy. To address these issues, we propose guiding the RL driving agent with a demonstration policy that need not be a highly optimized or expert-level controller. Specifically, we integrate a rule-based lane change controller with the Soft Actor Critic (SAC) algorithm to enhance exploration and learning efficiency. Our approach demonstrates improved driving performance and can be extended to other driving scenarios that can similarly benefit from demonstration-based guidance.

强化学习自动驾驶策略引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。