arXiv:2604.00372cs.CV2026-04

动态图神经网络自适应选择多模态局部特征,提升室内场景识别精度。

Dynamic Graph Neural Network with Adaptive Features Selection for RGB-D Based Indoor Scene Recognition

  • 构建动态图模型,基于注意力机制自适应筛选RGB与深度图关键局部特征。
  • 在SUN RGB-D和NYU Depth v2数据集上达到领先性能,显著优于现有方法。
  • 适合关注多模态融合与动态图建模的视觉识别研究者。

RGB-D多模态数据在室内场景识别中具有重要意义,其中深度图可描述场景的三维结构及物体间几何关系。以往研究指出,两种模态的局部特征对提升识别准确率至关重要,但如何自适应选择并有效利用这些关键局部特征仍是未解问题。本文提出一种动态图神经网络,通过自适应节点选择机制,从RGB和深度图中提取关键局部特征用于图建模。该模型将节点按三层次分组,表示物体间的远近关系,并根据注意力权重动态更新图结构。最终融合优化后的多模态特征实现场景识别。在SUN RGB-D和NYU Depth v2公开数据集上的实验表明,所提方法性能显著优于当前主流方法,验证了其在挖掘双模态关键特征方面的有效性。

原文摘要 · Abstract (English)

Multi-modality of color and depth, i.e., RGB-D, is of great importance in recent research of indoor scene recognition. In this kind of data representation, depth map is able to describe the 3D structure of scenes and geometric relations among objects. Previous works showed that local features of both modalities are vital for promotion of recognition accuracy. However, the problem of adaptive selection and effective exploitation on these key local features remains open in this field. In this paper, a dynamic graph model is proposed with adaptive node selection mechanism to solve the above problem. In this model, a dynamic graph is built up to model the relations among objects and scene, and a method of adaptive node selection is proposed to take key local features from both modalities of RGB and depth for graph modeling. After that, these nodes are grouped by three different levels, representing near or far relations among objects. Moreover, the graph model is updated dynamically according to attention weights. Finally, the updated and optimized features of RGB and depth modalities are fused together for indoor scene recognition. Experiments are performed on public datasets including SUN RGB-D and NYU Depth v2. Extensive results demonstrate that our method has superior performance when comparing to state-of-the-arts methods, and show that the proposed method is able to exploit crucial local features from both modalities of RGB and depth.

多模态图神经网络场景识别动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。