arXiv:2606.24767cs.CVcs.RO2026-06中稿 · RA-L 2026

用物体级结构化表示实现室内视觉重定位,提升语义理解与精度

Compact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual Relocalization

论文配图:Compact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual Relocalization
图 1 · 摘自论文原文
  • 基于多模态融合的开放词汇语义匹配,实现2D-3D物体精准对应
  • 引入基于DIOU的参考帧选择策略,支持可扩展场景下的定位
  • 设计双路2D-ICP损失优化姿态,显著提升定位稳定性与准确率

室内视觉重定位在新兴的空间智能与具身人工智能应用中至关重要。然而,以往研究多聚焦低层视觉方案,难以感知场景语义与结构,限制了可解释性与实用性。本文提出OpenReLoc系统,将场景中的语义、布局与几何信息组织为结构化物体级地图表示,仅使用物体单元驱动相机重定位。通过引入多模态机制融合开放词汇语义知识,实现高效的2D-3D物体匹配;设计面向物体的参考坐标系及基于距离-IoU(DIOU)的参考帧选择策略,支持大规模场景扩展;同时提出双路2D迭代最近像素损失,结合物体形状引导姿态优化。实验表明,OpenReLoc在多个数据集上均取得更优的重定位召回率与精度。代码将在接受后公开。

原文摘要 · Abstract (English)

Indoor visual relocalization plays a critical role in emerging spatial and embodied AI applications. However, prior research was predominantly devoted to low-level vision schemes, struggling to perceive scene semantics and compositions, which limits both interpretability and applicability. In this paper, we explore the issue of how to organize rich object information in a scene, including semantics, layout, and geometry, into a structured map representation, thereby utilizing object units exclusively to drive the camera relocalization task. To this end, we propose OpenReLoc, a camera relocalization system designed to provide scene understanding and accurate pose estimation capabilities. Leveraging recent foundation models, we first introduce a multi-modal mechanism to integrate open-vocabulary semantic knowledge for effective 2D-3D object matching. Additionally, we design object-oriented reference frames as position priors, paired with a reference frame selection strategy based on the Distance-IoU (DIOU), enabling extension to scalable scenes. Moreover, to ensure stable and accurate pose optimization, we also propose a dual-path 2D Iterative Closest Pixel loss guided by object shape. Experimental results demonstrate that OpenReLoc achieves superior relocalization recall and accuracy across various datasets. Our source code will be released upon acceptance.

视觉重定位物体级表示开放词汇姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。