arXiv:2512.22220cs.CVcs.AI2025-12

用视觉模型激活图定位被遮挡物品,提升机器人寻物效率。

On Extending Semantic Abstraction for Efficient Search of Hidden Objects

  • 利用视觉模型的响应图作为物体抽象表示
  • 首次尝试即准确找到3D隐藏物体位置,速度远超随机搜索
  • 适合家庭机器人寻物场景,依赖历史摆放数据优化搜索

语义抽象的核心观察是:2D 视觉语言模型(VLM)的相关性激活大致对应其对物体在场景中是否存在及位置的信心。因此,相关性图被视为‘抽象物体’的表示。我们在此框架下,学习针对隐藏物体(至少部分被遮挡、无法被 VLM 直接识别的物体)的 3D 定位与补全。该过程属于无结构搜索,可通过历史数据中物体常见放置位置实现更高效搜索。本模型能在首次尝试时准确识别隐藏物体的完整3D位置,显著快于朴素随机搜索。这些扩展希望为家用机器人提供节省时间与精力的寻物能力。

原文摘要 · Abstract (English)

Semantic Abstraction's key observation is that 2D VLMs' relevancy activations roughly correspond to their confidence of whether and where an object is in the scene. Thus, relevancy maps are treated as "abstract object" representations. We use this framework for learning 3D localization and completion for the exclusive domain of hidden objects, defined as objects that cannot be directly identified by a VLM because they are at least partially occluded. This process of localizing hidden objects is a form of unstructured search that can be performed more efficiently using historical data of where an object is frequently placed. Our model can accurately identify the complete 3D location of a hidden object on the first try significantly faster than a naive random search. These extensions to semantic abstraction hope to provide household robots with the skills necessary to save time and effort when looking for lost objects.

视觉定位隐藏物体机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。