arXiv:2506.13133cs.CV2025-06被引 3

用机器人感知约束优化视觉定位,小改动提升精度。

EmbodiedPlace: Learning Mixture-of-Features with Embodied Constraints for Visual Place Recognition

  • 基于机器人运动约束设计特征混合权重,动态融合多源信息
  • 仅增加25KB参数,每帧处理仅10微秒,提升0.9%识别率
  • 适合需要轻量化高精度定位的移动机器人应用

视觉场景识别(VPR)是计算机视觉中的场景导向图像检索问题,通常通过局部特征重排序来提升性能。在机器人领域,VPR也称为回环检测,强调序列中的时空一致性验证。然而,为VPR专门设计局部特征不切实际,依赖运动序列又存在局限。受此启发,我们提出一种新颖且简单的重排序方法,通过在具身约束下采用特征混合(MoF)方式精炼全局特征。首先,我们分析了具身约束在VPR中的可行性,并根据现有数据集将其分类为GPS标签、序列时间戳、局部特征匹配和自相似矩阵。随后,提出基于学习的MoF权重计算方法,使用多度量损失函数进行优化。实验表明,该方法在公共数据集上提升了当前最优性能,且额外计算开销极小。例如,仅增加25 KB参数,每帧处理耗时10微秒,就在Pitts-30k测试集上相较DINOv2基线提升了0.9%。

原文摘要 · Abstract (English)

Visual Place Recognition (VPR) is a scene-oriented image retrieval problem in computer vision in which re-ranking based on local features is commonly employed to improve performance. In robotics, VPR is also referred to as Loop Closure Detection, which emphasizes spatial-temporal verification within a sequence. However, designing local features specifically for VPR is impractical, and relying on motion sequences imposes limitations. Inspired by these observations, we propose a novel, simple re-ranking method that refines global features through a Mixture-of-Features (MoF) approach under embodied constraints. First, we analyze the practical feasibility of embodied constraints in VPR and categorize them according to existing datasets, which include GPS tags, sequential timestamps, local feature matching, and self-similarity matrices. We then propose a learning-based MoF weight-computation approach, utilizing a multi-metric loss function. Experiments demonstrate that our method improves the state-of-the-art (SOTA) performance on public datasets with minimal additional computational overhead. For instance, with only 25 KB of additional parameters and a processing time of 10 microseconds per frame, our method achieves a 0.9\% improvement over a DINOv2-based baseline performance on the Pitts-30k test set.

视觉定位具身智能特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。