arXiv:2506.18204cs.CVcs.AI2025-06中稿 · IEEE RAL被引 3

用傅里叶注意力融合多模态信息,提升低光噪声下的实时定位精度

Multimodal Fusion SLAM with Fourier Attention

  • 引入傅里叶自注意力与跨模态注意力,高效提取RGB和深度特征
  • 在TUM、TartanAir及自建数据集上实现比现有方法更优的鲁棒性
  • 已部署于安防机器人,结合GNSS-RTK与全局优化,支持实时运行

视觉SLAM在噪声干扰、光照变化和黑暗环境下面临挑战。基于学习的光流算法可利用多模态信息缓解此问题,但传统光流式视觉SLAM通常计算开销大。为此,本文提出FMF-SLAM,一种高效的多模态融合SLAM方法,采用快速傅里叶变换(FFT)提升算法效率。具体地,设计了一种基于傅里叶的自注意力与跨模态注意力机制,用于从RGB与深度信号中提取特征,并通过跨模态多尺度知识蒸馏增强多模态特征交互。同时,通过集成全球定位系统GNSS-RTK与全局束调整(Bundle Adjustment),验证了其在真实场景中的实用性与实时性能。在TUM、TartanAir及自建真实数据集上的实验表明,该方法在噪声、光照变化和暗光条件下均达到当前最优表现。代码与数据集已开源。

原文摘要 · Abstract (English)

Visual SLAM is particularly challenging in environments affected by noise, varying lighting conditions, and darkness. Learning-based optical flow algorithms can leverage multiple modalities to address these challenges, but traditional optical flow-based visual SLAM approaches often require significant computational resources.To overcome this limitation, we propose FMF-SLAM, an efficient multimodal fusion SLAM method that utilizes fast Fourier transform (FFT) to enhance the algorithm efficiency. Specifically, we introduce a novel Fourier-based self-attention and cross-attention mechanism to extract features from RGB and depth signals. We further enhance the interaction of multimodal features by incorporating multi-scale knowledge distillation across modalities. We also demonstrate the practical feasibility of FMF-SLAM in real-world scenarios with real time performance by integrating it with a security robot by fusing with a global positioning module GNSS-RTK and global Bundle Adjustment. Our approach is validated using video sequences from TUM, TartanAir, and our real-world datasets, showcasing state-of-the-art performance under noisy, varying lighting, and dark conditions.Our code and datasets are available at https://github.com/youjie-zhou/FMF-SLAM.git.

SLAM多模态融合傅里叶注意力实时定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。