用稀疏时空融合提升视觉-激光雷达里程计精度与鲁棒性
DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion
- 通过稀疏激光点查询实现多模态数据高效融合
- 时间交互模块减少累积误差,长序列定位更稳定
- 适合需要实时高精度定位的自动驾驶系统
视觉-激光雷达里程计是自主系统定位的关键组件,但实现高精度与强鲁棒性仍具挑战。传统方法常受传感器错位影响,未能充分挖掘时序信息,且需大量人工调参适配不同传感器配置。为此,我们提出DVLO4D,一种新型视觉-激光雷达里程计框架,通过稀疏时空融合提升性能。核心创新包括:(1) 稀疏查询融合,利用稀疏激光点查询实现高效多模态融合;(2) 时间交互与更新模块,将时序预测位置与当前帧数据结合,提供更优姿态估计初值,增强对累积误差的鲁棒性;(3) 时间片段训练策略与集体平均损失机制,跨多帧聚合损失,实现全局优化,显著降低长序列尺度漂移。在KITTI和Argoverse Odometry数据集上的实验表明,DVLO4D在姿态精度与鲁棒性上均达到当前最优水平。此外,推理时间仅82毫秒,具备实时部署潜力。
原文摘要 · Abstract (English)
Visual-LiDAR odometry is a critical component for autonomous system localization, yet achieving high accuracy and strong robustness remains a challenge. Traditional approaches commonly struggle with sensor misalignment, fail to fully leverage temporal information, and require extensive manual tuning to handle diverse sensor configurations. To address these problems, we introduce DVLO4D, a novel visual-LiDAR odometry framework that leverages sparse spatial-temporal fusion to enhance accuracy and robustness. Our approach proposes three key innovations: (1) Sparse Query Fusion, which utilizes sparse LiDAR queries for effective multi-modal data fusion; (2) a Temporal Interaction and Update module that integrates temporally-predicted positions with current frame data, providing better initialization values for pose estimation and enhancing model's robustness against accumulative errors; and (3) a Temporal Clip Training strategy combined with a Collective Average Loss mechanism that aggregates losses across multiple frames, enabling global optimization and reducing the scale drift over long sequences. Extensive experiments on the KITTI and Argoverse Odometry dataset demonstrate the superiority of our proposed DVLO4D, which achieves state-of-the-art performance in terms of both pose accuracy and robustness. Additionally, our method has high efficiency, with an inference time of 82 ms, possessing the potential for the real-time deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。