提出可学习的相机与激光雷达融合方法,提升自动驾驶端到端感知与预测性能。
Li-ViP3D++: Query-Gated Deformable Camera-LiDAR Fusion for End-to-End Perception and Trajectory Prediction
- 在查询空间中通过可学习门控实现多视角图像与激光雷达特征融合。
- 在nuScenes数据集上检测准确率(mAP)达0.502,误报率降低至0.147。
- 适合关注自动驾驶多模态感知与轨迹预测的研究者和工程师。
从原始传感器数据实现端到端感知与轨迹预测是自动驾驶的关键能力。传统模块化流程限制信息流动并放大上游误差。近期基于查询的全可微感知与预测(PnP)模型缓解了这些问题,但摄像头与激光雷达在查询空间中的互补性尚未充分探索。现有模型常依赖启发式对齐和离散选择步骤,难以充分利用信息并可能引入偏差。本文提出Li-ViP3D++,一种基于查询的多模态端到端PnP框架,引入查询门控可变形融合(QGDF),在查询空间内融合多视角RGB图像与激光雷达数据。QGDF(i)通过掩码注意力聚合跨相机与特征层级的图像证据,(ii)通过可学习每查询偏移的全可微鸟瞰图采样提取激光雷达上下文,(iii)采用查询条件门控自适应加权视觉与几何线索。该架构在单一端到端模型中联合优化检测、跟踪与多假设轨迹预测。在nuScenes数据集上,Li-ViP3D++提升端到端行为与检测质量,达到更高EPA(0.335)与mAP(0.502),同时显著降低误报率(FP ratio 0.147),且比前代Li-ViP3D更快(139.82 ms vs. 145.91 ms)。结果表明,查询空间全可微融合能增强端到端PnP鲁棒性而不牺牲部署性。
原文摘要 · Abstract (English)
End-to-end perception and trajectory prediction from raw sensor data is one of the key capabilities for autonomous driving. Modular pipelines restrict information flow and can amplify upstream errors. Recent query-based, fully differentiable perception-and-prediction (PnP) models mitigate these issues, yet the complementarity of cameras and LiDAR in the query-space has not been sufficiently explored. Models often rely on fusion schemes that introduce heuristic alignment and discrete selection steps which prevent full utilization of available information and can introduce unwanted bias. We propose Li-ViP3D++, a query-based multimodal PnP framework that introduces Query-Gated Deformable Fusion (QGDF) to integrate multi-view RGB and LiDAR in query space. QGDF (i) aggregates image evidence via masked attention across cameras and feature levels, (ii) extracts LiDAR context through fully differentiable BEV sampling with learned per-query offsets, and (iii) applies query-conditioned gating to adaptively weight visual and geometric cues per agent. The resulting architecture jointly optimizes detection, tracking, and multi-hypothesis trajectory forecasting in a single end-to-end model. On nuScenes, Li-ViP3D++ improves end-to-end behavior and detection quality, achieving higher EPA (0.335) and mAP (0.502) while substantially reducing false positives (FP ratio 0.147), and it is faster than the prior Li-ViP3D variant (139.82 ms vs. 145.91 ms). These results indicate that query-space, fully differentiable camera-LiDAR fusion can increase robustness of end-to-end PnP without sacrificing deployability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。