arXiv:2510.26131cs.CVcs.RO2025-10被引 1

用注意力机制提升RGB-D SLAM的帧关联,增强室内定位精度

Exploring Object-Aware Attention Guided Frame Association for RGB-D SLAM

  • 通过网络梯度提取层间注意力信息,融合到CNN特征中
  • 在大场景下帧关联准确率显著优于基线方法
  • 适合需要高精度定位的室内SLAM应用

注意力模型在多个领域展现出强大能力。可视化技术如类别激活图可揭示卷积神经网络(CNN)的推理过程,利用网络梯度能识别图像识别任务中网络关注的区域。这些梯度还可与CNN特征结合,定位更通用、任务相关的显著区域。然而,将基于梯度的注意力信息直接融入CNN表示以实现语义物体理解仍较少见。该融合对同时定位与建图(SLAM)等视觉任务尤为有益,因为包含空间注意力目标位置的CNN表征可提升性能。本文提出将任务特定的网络注意力用于RGB-D室内SLAM,具体为将由网络梯度导出的层间注意力信息与CNN特征表示融合,以改进帧关联性能。实验结果表明,相较于基线方法,本方法在大环境下的表现显著提升。

原文摘要 · Abstract (English)

Attention models have recently emerged as a powerful approach, demonstrating significant progress in various fields. Visualization techniques, such as class activation mapping, provide visual insights into the reasoning of convolutional neural networks (CNNs). Using network gradients, it is possible to identify regions where the network pays attention during image recognition tasks. Furthermore, these gradients can be combined with CNN features to localize more generalizable, task-specific attentive (salient) regions within scenes. However, explicit use of this gradient-based attention information integrated directly into CNN representations for semantic object understanding remains limited. Such integration is particularly beneficial for visual tasks like simultaneous localization and mapping (SLAM), where CNN representations enriched with spatially attentive object locations can enhance performance. In this work, we propose utilizing task-specific network attention for RGB-D indoor SLAM. Specifically, we integrate layer-wise attention information derived from network gradients with CNN feature representations to improve frame association performance. Experimental results indicate improved performance compared to baseline methods, particularly for large environments.

SLAM注意力机制深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。