arXiv:2606.04884cs.RO2026-06

让自动驾驶模型同时学会多种驾驶风格,还能自由切换。

D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving

论文配图:D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving
图 1 · 摘自论文原文
  • 用双解耦扩散机制分离风格与物理动作,分别生成候选轨迹和独立预测路径。
  • 在NAVSIM基准上达到91.3 PDMS和87.5 EPDMS的最优性能,支持多风格选择。
  • 适合需要高可解释性与风格可控性的自动驾驶系统研发人员。

传统端到端自动驾驶框架在训练时因人类示范数据差异大,常出现“风格平均化”问题,导致策略同质、不可控且存在运动学安全隐患。为此,本文提出D³-MoE(双解耦扩散混合专家模型),沿行为与物理两个轴解耦轨迹建模:在行为轴上,通过风格条件扩散过程并行生成多风格候选轨迹,由下游模块根据偏好或评分选择最优;在物理轴上,纵向与横向路由模块在推理时独立激活对应专家,无需人工标注,仅依赖正交真值运动学进行自监督训练。这些专家以扩散Transformer架构实现,配备风格条件AdaLN和非对称侧向融合交叉注意力,分别预测对应物理状态后重构为连贯轨迹。在挑战性NAVSIM基准上的实验证明,D³-MoE默认表现达88.2 PDMS与84.3 EPDMS,采用最佳三选一集成策略后提升至91.3 PDMS与87.5 EPDMS。定量与定性分析共同验证其在规划质量与风格可控性方面的优势。

原文摘要 · Abstract (English)

Traditional end-to-end autonomous driving frameworks frequently suffer from the "style-averaging" dilemma when trained on high-variance human demonstrations, yielding homogenized, style-uncontrollable, and even kinematically unsafe policies. To overcome this limitation, we present D$^3$-MoE (Dual Disentangled Diffusion Mixture-of-Experts), which disentangles trajectory modeling along two complementary axes. On the behavioral axis, generation is decoupled from selection: a style-conditioned diffusion process synthesizes multi-style candidate trajectories in parallel within a single scene, allowing a downstream module to select the optimal trajectory based on user preference or an evaluation score. On the physical axis, decoupled longitudinal and lateral routers activate their respective experts during inference time, trained without manual labels using self-supervised targets from orthogonal ground-truth kinematics. These activated experts, architected as Diffusion Transformers (DiT) and equipped with style-conditioned AdaLN and asymmetric lateral-fusion cross-attention, independently predict their corresponding physical state before being reassembled into a unified, kinematically coherent trajectory. Extensive evaluations on the challenging NAVSIM benchmark demonstrate that D$^3$-MoE achieves state-of-the-art planning performance, reaching 88.2 PDMS and 84.3 EPDMS by default. Moreover, our Best-of-Three ensemble strategy effectively broadens the multi-modal solution space, raising performance to 91.3 PDMS and 87.5 EPDMS. Both quantitative and qualitative analyses jointly confirm the framework's advantages in planning quality and style controllability.

自动驾驶扩散模型风格控制Mixture-of-Experts

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。