arXiv:2511.19221cs.CVcs.RO2025-11被引 9

将2D/3D感知融合进视觉语言模型,提升自动驾驶在长尾场景下的鲁棒性。

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

  • 用世界坐标与置信度编码统一2D/3D感知任务
  • 在COCO和nuScenes上分别达51.7/58.9 mAP
  • 直接输出感知结果与轨迹控制,适合端到端系统

自动驾驶严重依赖精准稳定的空间感知,但现有视觉语言模型在空间定位与理解上能力不足,导致系统在长尾场景和复杂交互中易失效。为此,我们提出Percept-WAM,首个在单一视觉语言模型中隐式融合2D/3D场景理解能力的感知增强型世界意识-动作模型。该模型将2D/3D感知任务统一为世界-图像坐标(World-PV)与世界-俯视图(World-BEV)令牌,同时编码空间坐标与置信度。通过网格条件预测机制,结合IoU感知评分与并行自回归解码,显著提升远距离、小目标及长尾场景下的感知稳定性。此外,利用预训练视觉语言模型参数保留通用智能(如逻辑推理),可直接输出感知结果与轨迹控制信号。实验表明,Percept-WAM在下游感知基准上达到或超越经典检测器与分割器,在COCO 2D检测与nuScenes BEV 3D检测上分别取得51.7与58.9 mAP。集成轨迹解码器后,在nuScenes与NAVSIM上进一步提升规划性能,例如在NAVSIM上比DiffusionDrive高出2.1的PMDS。定性结果也验证其出色的开放词汇与长尾泛化能力。

原文摘要 · Abstract (English)

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak at spatial grounding and understanding, and VLA systems built on them therefore show limited perception and localization ability. To address these challenges, we introduce Percept-WAM, a perception-enhanced World-Awareness-Action Model that is the first to implicitly integrate 2D/3D scene understanding abilities within a single vision-language model (VLM). Instead of relying on QA-style spatial reasoning, Percept-WAM unifies 2D/3D perception tasks into World-PV and World-BEV tokens, which encode both spatial coordinates and confidence. We propose a grid-conditioned prediction mechanism for dense object perception, incorporating IoU-aware scoring and parallel autoregressive decoding, improving stability in long-tail, far-range, and small-object scenarios. Additionally, Percept-WAM leverages pretrained VLM parameters to retain general intelligence (e.g., logical reasoning) and can output perception results and trajectory control outputs directly. Experiments show that Percept-WAM matches or surpasses classical detectors and segmenters on downstream perception benchmarks, achieving 51.7/58.9 mAP on COCO 2D detection and nuScenes BEV 3D detection. When integrated with trajectory decoders, it further improves planning performance on nuScenes and NAVSIM, e.g., surpassing DiffusionDrive by 2.1 in PMDS on NAVSIM. Qualitative results further highlight its strong open-vocabulary and long-tail generalization.

自动驾驶视觉语言模型多模态感知端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。