arXiv:2503.01130cs.CV2025-03CVPR被引 3

AirRoom通过融合物体信息提升室内场景重识别准确率

AirRoom: Objects Matter in Room Reidentification

  • 从全局到物体区域、分割、关键点,分层利用物体信息
  • 在4个新数据集上比顶尖模型最高提升80%
  • 模块可替换,对视角变化鲁棒,适合机器人导航

室场景重识别(ReID)是增强现实和家庭护理机器人等领域的关键挑战任务。现有视觉场景识别方法多依赖全局描述符或聚合局部特征,在密集布满人工物体的复杂室内环境中表现不佳,常忽略物体导向信息。为此,我们提出AirRoom,一种基于物体感知的端到端流程,融合多层次物体信息——从全局上下文到物体块、物体分割与关键点,并采用粗到精的检索策略。在四个新构建的数据集(MPReID、HMReID、GibsonReID、ReplicaReID)上的大量实验表明,AirRoom在几乎所有评估指标上均优于当前最优模型,性能提升达6%至80%。此外,该框架具有显著灵活性,各模块可自由替换而不影响整体效果,且在多种视角变化下仍保持稳健一致的表现。

原文摘要 · Abstract (English)

Room reidentification (ReID) is a challenging yet essential task with numerous applications in fields such as augmented reality (AR) and homecare robotics. Existing visual place recognition (VPR) methods, which typically rely on global descriptors or aggregate local features, often struggle in cluttered indoor environments densely populated with man-made objects. These methods tend to overlook the crucial role of object-oriented information. To address this, we propose AirRoom, an object-aware pipeline that integrates multi-level object-oriented information-from global context to object patches, object segmentation, and keypoints-utilizing a coarse-to-fine retrieval approach. Extensive experiments on four newly constructed datasets-MPReID, HMReID, GibsonReID, and ReplicaReID-demonstrate that AirRoom outperforms state-of-the-art (SOTA) models across nearly all evaluation metrics, with improvements ranging from 6% to 80%. Moreover, AirRoom exhibits significant flexibility, allowing various modules within the pipeline to be substituted with different alternatives without compromising overall performance. It also shows robust and consistent performance under diverse viewpoint variations.

场景识别物体感知机器人导航多尺度特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。