arXiv:2508.16947cs.ROcs.AI2025-08

让自动驾驶懂用户偏好,生成个性化驾驶轨迹。

Drive As You Like: Multi-Head Diffusion with Reinforcement Learning for Personalized Driving

  • 用多头扩散模型+强化学习,动态理解用户意图。
  • 在nuPlan上实测,轨迹质量高且能实时生成。
  • 适合需要个性化的智能驾驶系统研发者。

尽管进展显著,基于模仿学习的自动驾驶规划器仍主要复现高频偏差行为,忽视了人类驾驶的固有多样性。现有系统难以从人机交互和环境上下文中理解用户意图。在真实场景部署中,运动规划需适应多样、情境依赖的用户偏好,以支持多元驾驶服务,这要求能够解析用户意图并相应调整行为。然而,现有方法缺乏此类以用户为中心的能力,既未显式建模用户意图,也无法灵活调整策略。为此,我们提出一种由强化学习引导的多策略框架,结合基于扩散模型的多头规划器(M-Diffusion Planner)与基于大语言模型的语义理解,实现对用户意图的动态感知和生成多样化、偏好一致的轨迹。为平衡轨迹质量和策略一致性,采用两阶段训练范式:首先通过模仿学习确保每个策略头达到安全且高质量的规划;其次通过受限的组相对策略优化(GRPO)进一步对齐各头与用户偏好。在nuPlan基准测试中,无论开环还是闭环设置下,实验均表明该方法性能具有竞争力,满足实时规划需求,并有效对齐用户意图。

原文摘要 · Abstract (English)

Despite significant progress, imitation learning-based autonomous driving planners remain largely restricted to reproducing high-frequency biased behaviors, overlooking the inherent behavioral diversity of human driving. Moreover, existing systems struggle to understand user intent from human interactions and environmental contexts. In real-world advanced deployment, motion planning must accommodate diverse, context-dependent user preferences to support heterogeneous driving services, requiring the ability to interpret human intent and adapt behavior accordingly. However, existing approaches lack such user-oriented capabilities, as they neither explicitly model user intent nor enable flexible policy adaptation. To bridge this gap, we propose an RL-guided multi-strategy framework with a diffusion-based multi-head planner(M-Diffusion Planner) integrated with LLM-based semantic understanding, enabling dynamic perception of user intent and generation of diverse, preference-aligned trajectories. To balance trajectory quality and strategy alignment, we adopt a two-stage training paradigm: first, imitation learning ensures each policy head achieves safe and high-quality planning; second, constrained Group Relative Policy Optimization (GRPO) further aligns each head with user preferences. Experiments on the nuPlan benchmark, under both open-loop and closed-loop settings, demonstrate competitive performance while meeting real-time planning requirements and effectively aligning with user intent.

自动驾驶扩散模型强化学习个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。