arXiv:2512.07135cs.CVcs.AI2025-12被引 1

用专家混合与强化学习,让自动驾驶更懂场景地规划路径。

TrajMoE: Scene-Adaptive Trajectory Planning with Mixture of Experts and Reinforcement Learning

  • 用专家混合模型按场景动态选择轨迹先验
  • 强化学习优化轨迹评分,提升规划准确性
  • 适配多感知结构,适合复杂城市驾驶场景

当前自动驾驶系统多采用端到端框架,直接从图像等传感器输入映射到轨迹空间。已有研究表明,提供合理的轨迹先验可提升规划性能。但现有方法忽视两点:1)不同驾驶场景需不同轨迹先验;2)轨迹评估机制缺乏策略驱动的迭代优化,受限于单阶段监督训练。为此,我们从两方面改进:针对问题1,采用混合专家(MoE)模型根据场景自适应选择轨迹先验;针对问题2,利用强化学习对轨迹评分机制进行精细化调整。此外,融合多种感知骨干网络以增强感知特征表达。所提模型在navsim ICCV基准上获得51.08分,位列第三。

原文摘要 · Abstract (English)

Current autonomous driving systems often favor end-to-end frameworks, which take sensor inputs like images and learn to map them into trajectory space via neural networks. Previous work has demonstrated that models can achieve better planning performance when provided with a prior distribution of possible trajectories. However, these approaches often overlook two critical aspects: 1) The appropriate trajectory prior can vary significantly across different driving scenarios. 2) Their trajectory evaluation mechanism lacks policy-driven refinement, remaining constrained by the limitations of one-stage supervised training. To address these issues, we explore improvements in two key areas. For problem 1, we employ MoE to apply different trajectory priors tailored to different scenarios. For problem 2, we utilize Reinforcement Learning to fine-tune the trajectory scoring mechanism. Additionally, we integrate models with different perception backbones to enhance perceptual features. Our integrated model achieved a score of 51.08 on the navsim ICCV benchmark, securing third place.

自动驾驶轨迹规划MoE强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。