arXiv:2603.08544cs.ROcs.LG2026-03

让机器人通过场景特征关系智能找东西,效率比现有方法高20%。

The Neural Compass: Probabilistic Relative Feature Fields for Robotic Search

  • 用无标签数据训练特征场模型,自动学习物体间的空间关联
  • 在Matterport3D中搜索任务上比最强基线快20%,达人类80%表现
  • 适合做视觉导航与自主搜索的机器人研究者

物体共现是高效寻找未知环境中物体的关键线索。通常人们会在厨房找杯子,并将冰箱作为身处厨房的证据。这类先验知识也被用于人工智能代理,但多依赖标注数据或语言模型查询。目前尚不清楚这些关系能否仅从无标签观测中隐式学习。本文提出ProReFF,一种基于预训练视觉语言模型特征的相对特征场模型,可预测特征间的相对分布。同时引入一种学习策略,能从无标签且可能矛盾的数据中对齐不一致观测,形成一致的相对分布。针对下游物体搜索任务,设计了一种利用预测特征分布作为语义先验的智能体,引导探索至更可能包含目标的区域。大量实验表明,ProReFF能有效捕捉自然场景中的有意义相对特征分布,并验证了对齐步骤的有效性。在Matterport3D模拟器的100个挑战中,该智能体比特征基线高出20%效率,最高达到人类80%的表现。

原文摘要 · Abstract (English)

Object co-occurrences provide a key cue for finding objects successfully and efficiently in unfamiliar environments. Typically, one looks for cups in kitchens and views fridges as evidence of being in a kitchen. Such priors have also been exploited in artificial agents, but they are typically learned from explicitly labeled data or queried from language models. It is still unclear whether these relations can be learned implicitly from unlabeled observations alone. In this work, we address this problem and propose ProReFF, a feature field model trained to predict relative distributions of features obtained from pre-trained vision language models. In addition, we introduce a learning-based strategy that enables training from unlabeled and potentially contradictory data by aligning inconsistent observations into a coherent relative distribution. For the downstream object search task, we propose an agent that leverages predicted feature distributions as a semantic prior to guide exploration toward regions with a high likelihood of containing the object. We present extensive evaluations demonstrating that ProReFF captures meaningful relative feature distributions in natural scenes and provides insight into the impact of our proposed alignment step. We further evaluate the performance of our search agent in 100 challenges in the Matterport3D simulator, comparing with feature-based baselines and human participants. The proposed agent is 20% more efficient than the strongest baseline and achieves up to 80% of human performance.

机器人搜索视觉定位特征场无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。