提升机器人感知中多传感器融合的时序一致性,减少定位误差。
LiDAR-BIND-T: Improved and Temporally Consistent Sensor Modality Translation and Fusion for Robotic Applications
- 引入时序嵌入相似性与运动对齐损失,保持连续帧间特征一致。
- 在Cartographer SLAM中实现更低轨迹误差和更高占用图精度。
- 适合需要稳定感知的自动驾驶与机器人导航场景。
本文扩展了LiDAR-BIND框架,该框架将异构传感器(雷达、声呐)绑定至由激光雷达定义的潜在空间,并引入显式时序一致性机制。提出三项改进:(i) 时序嵌入相似性,对齐连续潜在表示;(ii) 运动对齐变换损失,匹配预测与真实激光雷达间的位移;(iii) 基于专用时序模块的窗口化时序融合。同时优化模型结构以更好保留空间结构。在雷达/声呐到激光雷达的转换任务上,实验显示时序与空间一致性显著提升,使基于Cartographer的SLAM系统轨迹误差更低、占用图精度更高。提出基于Fréchet视频运动距离(FVMD)和相关峰值距离的新度量指标,用于评估时序质量。所提出的时序LiDAR-BIND(LiDAR-BIND-T)在保持模块化融合的同时大幅增强时序稳定性,显著提升下游SLAM的鲁棒性与性能。
原文摘要 · Abstract (English)
This paper extends LiDAR-BIND, a modular multi-modal fusion framework that binds heterogeneous sensors (radar, sonar) to a LiDAR-defined latent space, with mechanisms that explicitly enforce temporal consistency. We introduce three contributions: (i) temporal embedding similarity that aligns consecutive latent representations, (ii) a motion-aligned transformation loss that matches displacement between predictions and ground truth LiDAR, and (iii) windowed temporal fusion using a specialised temporal module. We further update the model architecture to better preserve spatial structure. Evaluations on radar/sonar-to-LiDAR translation demonstrate improved temporal and spatial coherence, yielding lower absolute trajectory error and better occupancy map accuracy in Cartographer-based SLAM (Simultaneous Localisation and Mapping). We propose different metrics based on the Fréchet Video Motion Distance (FVMD) and a correlation-peak distance metric providing practical temporal quality indicators to evaluate SLAM performance. The proposed temporal LiDAR-BIND, or LiDAR-BIND-T, maintains modular modality fusion while substantially enhancing temporal stability, resulting in improved robustness and performance for downstream SLAM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。