arXiv:2507.19146cs.ROcs.LG2025-07中稿 · IEEE/RSJ Internati…

用师生框架自动生成复杂交通行为,提升自动驾驶训练效果

Diverse and Adaptive Behavior Curriculum for Autonomous Driving: A Student-Teacher Framework with Multi-Agent RL

  • 教师用图结构多智能体RL动态生成不同难度交通行为
  • 学生在自适应课程中训练,驾驶表现优于规则场景训练
  • 适合需要真实复杂路况训练的自动驾驶研发团队

自动驾驶在复杂真实交通环境中面临挑战,需安全应对常见与关键场景。强化学习(RL)虽能在仿真中通过试错学习,但当前训练依赖规则化交通场景,限制泛化能力。现有场景生成方法过度关注关键场景,忽视日常驾驶行为的平衡。课程学习通过逐步增加任务难度,可提升RL驾驶策略的鲁棒性与覆盖度。然而,现有研究多依赖人工设计课程,侧重场景布局而非交通行为动态。本文提出一种新型师生框架实现自动课程学习:教师为基于图的多智能体RL组件,自适应生成多样难度的交通行为;学生为具备部分可观测性的深度RL代理,反映真实感知约束。通过性能反馈动态调整任务难度,确保涵盖从常规到关键行为的全范围训练。实验表明,教师能生成多样化交通行为;学生在自动课程下训练后,奖励更高,驾驶行为更均衡且主动。

原文摘要 · Abstract (English)

Autonomous driving faces challenges in navigating complex real-world traffic, requiring safe handling of both common and critical scenarios. Reinforcement learning (RL), a prominent method in end-to-end driving, enables agents to learn through trial and error in simulation. However, RL training often relies on rule-based traffic scenarios, limiting generalization. Additionally, current scenario generation methods focus heavily on critical scenarios, neglecting a balance with routine driving behaviors. Curriculum learning, which progressively trains agents on increasingly complex tasks, is a promising approach to improving the robustness and coverage of RL driving policies. However, existing research mainly emphasizes manually designed curricula, focusing on scenery and actor placement rather than traffic behavior dynamics. This work introduces a novel student-teacher framework for automatic curriculum learning. The teacher, a graph-based multi-agent RL component, adaptively generates traffic behaviors across diverse difficulty levels. An adaptive mechanism adjusts task difficulty based on student performance, ensuring exposure to behaviors ranging from common to critical. The student, though exchangeable, is realized as a deep RL agent with partial observability, reflecting real-world perception constraints. Results demonstrate the teacher's ability to generate diverse traffic behaviors. The student, trained with automatic curricula, outperformed agents trained on rule-based traffic, achieving higher rewards and exhibiting balanced, assertive driving.

自动驾驶强化学习课程学习多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。