通过细粒度特征与场景图增强,实现高效图像目标导航。
Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement
- 用空间-通道注意力融合目标与观测特征
- 在Gibson和HM3D上达领先性能,推理速度53.5帧/秒
- 适合实时端到端导航应用,模型已开源
本文提出一种新型图像目标导航方法RFSG,聚焦于在有限图像数据下挖掘目标、观测与环境间的细粒度关联,同时保持导航架构简洁轻量。为此,我们设计了空间-通道注意力机制,使网络能学习多维特征的重要性,融合目标与观测特征;引入自蒸馏机制进一步提升特征表达能力。考虑到导航需依赖周围环境信息以提高效率,我们构建图像场景图,在图像与物体层面建立特征关联,有效编码周边场景信息。在Gibson和HM3D数据集上进行跨场景性能验证,所提方法在主流方法中达到最先进水平,且在RTX3080上实现高达53.5帧/秒的推理速度,推动真实场景中端到端图像目标导航的实现。代码与模型已公开于:https://github.com/nubot-nudt/RFSG。
原文摘要 · Abstract (English)
In this paper, we introduce a novel image-goal navigation approach, named RFSG. Our focus lies in leveraging the fine-grained connections between goals, observations, and the environment within limited image data, all the while keeping the navigation architecture simple and lightweight. To this end, we propose the spatial-channel attention mechanism, enabling the network to learn the importance of multi-dimensional features to fuse the goal and observation features. In addition, a selfdistillation mechanism is incorporated to further enhance the feature representation capabilities. Given that the navigation task needs surrounding environmental information for more efficient navigation, we propose an image scene graph to establish feature associations at both the image and object levels, effectively encoding the surrounding scene information. Crossscene performance validation was conducted on the Gibson and HM3D datasets, and the proposed method achieved stateof-the-art results among mainstream methods, with a speed of up to 53.5 frames per second on an RTX3080. This contributes to the realization of end-to-end image-goal navigation in realworld scenarios. The implementation and model of our method have been released at: https://github.com/nubot-nudt/RFSG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。