arXiv:2509.18609cs.RO2025-09被引 1

PIE通过融合感知与交互建模,实现自动驾驶端到端路径规划

PIE: Perception and Interaction Enhanced End-to-End Motion Planning for Autonomous Driving

  • 引入双向Mamba融合处理多模态输入,减少信息损失
  • 在NAVSIM上达88.9的PDM分数,超越现有最先进方法
  • 适合研究自动驾驶决策与轨迹生成的工程师和研究员

端到端运动规划有望简化复杂的自动驾驶系统流程,但场景理解与决策预测仍面临挑战。本文提出PIE框架,融合先进感知、推理与意图建模,动态捕捉自车与周围交通参与者之间的交互。该框架采用双向Mamba融合机制,缓解摄像头与激光雷达输入在多模态融合中的数据压缩损失;并设计基于Mamba与专家混合(Mixture-of-Experts)的增强型解码器,支持场景适配的锚点选择与自适应轨迹推断。此外,引入动作-运动交互模块,利用周围车辆状态预测优化自车规划。在NAVSIM基准测试中,PIE未使用集成或数据增强,即取得88.9的PDM得分和85.6的EPDM得分,优于现有最先进方法。定量与定性分析表明,PIE可稳定生成可行且高质量的自车轨迹。

原文摘要 · Abstract (English)

End-to-end motion planning is promising for simplifying complex autonomous driving pipelines. However, challenges such as scene understanding and effective prediction for decision-making continue to present substantial obstacles to its large-scale deployment. In this paper, we present PIE, a pioneering framework that integrates advanced perception, reasoning, and intention modeling to dynamically capture interactions between the ego vehicle and surrounding agents. It incorporates a bidirectional Mamba fusion that addresses data compression losses in multimodal fusion of camera and LiDAR inputs, alongside a novel reasoning-enhanced decoder integrating Mamba and Mixture-of-Experts to facilitate scene-compliant anchor selection and optimize adaptive trajectory inference. PIE adopts an action-motion interaction module to effectively utilize state predictions of surrounding agents to refine ego planning. The proposed framework is thoroughly validated on the NAVSIM benchmark. PIE, without using any ensemble and data augmentation techniques, achieves an 88.9 PDM score and 85.6 EPDM score, surpassing the performance of prior state-of-the-art methods. Comprehensive quantitative and qualitative analyses demonstrate that PIE is capable of reliably generating feasible and high-quality ego trajectories.

自动驾驶端到端规划多模态融合轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。