arXiv:2607.20175cs.CV2026-07

用自适应路由增强感知先验,实现端到端自动驾驶的高效规划。

PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving

论文配图:PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving
图 1 · 摘自论文原文
  • 引入可学习的专家路由机制,动态融合感知先验与观测特征。
  • 在NAVSIM v1和v2上分别达90.4和90.2的PDMS得分,性能领先。
  • 无需测试时搜索或重排序,单摄像头输入即可生成轨迹。

冻结的感知基础模型蕴含丰富的几何、语义与动态知识,但其有限的条件接口可能削弱任务相关线索,静态融合无法按场景调整专家贡献。本文将此问题视为先验到规划的迁移挑战,提出PerceptDrive:一种基于自适应专家路由的感知先验世界-动作建模框架。该框架输入来自驾驶适配型提供者的教师蒸馏先验,以及来自冻结自监督视频编码器的密集观测隐向量,通过专家特异性查询分支处理信号,并以先验保留目标锚定各分支。路由器基于共享场景表示预测软门控,整合专家条件后生成轨迹。训练中,使用特权规则基子指标估计分支级轨迹草图作为软门控蒸馏目标。预测的动作自由未来隐状态条件流匹配动作器。推理时,特权组件被移除;仅需单前视摄像头,在每个规划步骤生成一条轨迹,无需测试时评分、重排序或搜索。实验表明,PerceptDrive在NAVSIM v1上达到90.4 PDMS,在NAVSIM v2上达到90.2 EPDMS,优于现有方法。消融实验验证了先验保留与场景条件路由的互补增益,以及对三类先验的不同依赖程度。结果证明,保持并自适应路由感知先验,可在无测试时候选选择的情况下提升直接规划性能。

原文摘要 · Abstract (English)

Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion cannot adjust expert contributions to each scene. We cast this challenge as the prior-to-plan transfer problem and introduce PerceptDrive, a perception prior world-action modeling framework with adaptive expert routing. PerceptDrive feeds teacher-distilled priors from a frozen, driving-adapted provider and dense observation latents from a frozen self-supervised video encoder into a trainable expert-routed world-action model. Expert-specific query branches process these signals, while a prior-retention objective anchors each branch to its prior. A router predicts soft gates from a shared scene representation and combines the expert conditions before trajectory generation. During training, privileged rule-based sub-metric estimates for branch-specific trajectory drafts provide soft-gate distillation targets. The predicted action-free future latent conditions a flow-matching actor. At inference, privileged components are absent; with one front-facing camera, PerceptDrive generates one trajectory per planning step without test-time scoring, reranking, or search. Experiments show that PerceptDrive achieves state-of-the-art performance with 90.4 PDMS on NAVSIM v1 and 90.2 EPDMS on NAVSIM v2, outperforming existing methods. Ablations confirm complementary gains from prior retention and scene-conditioned routing, alongside differential reliance on the three priors. These results demonstrate that preserving and adaptively routing perception priors improves direct planning without test-time candidate selection.

自动驾驶感知先验专家路由端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。