arXiv:2605.14704cs.CVcs.AI2026-05

让模型学会推理被遮挡物体的位置,提升视觉理解能力。

SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization

论文配图:SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
图 1 · 摘自论文原文
  • 构建2D空间推理任务,基于场景上下文推断不可见物体位置。
  • 最强模型仅达15.2%准确率,说明隐含区域推理仍困难。
  • 适合关注常识推理与空间认知的视觉语言模型研究者。

在真实场景中,目标物体可能位于不可见区域。人类常能通过上下文和常识推断被遮挡物体的位置,但当前视觉语言模型(VLMs)仍难以胜任此任务。为此,我们提出SceneFunRI基准,基于SceneFun3D数据集,通过半自动流程将任务转化为2D空间推理问题,包含855个实例,要求模型根据任务指令和常识推理定位不可见的功能性物体。最强基线模型Gemini 3 Flash仅达到CAcc@75为15.20、mIoU为0.74、Dist为28.65。我们将其提示策略分为三类:强指令提示、基于推理的提示和空间排除法(SPoE),结果表明当前VLM在不可见区域推理上仍不稳定,亟需更紧密融合任务意图、常识先验、空间定位与不确定性感知搜索的模型。

原文摘要 · Abstract (English)

In real-world scenes, target objects may reside in regions that are not visible. While humans can often infer the locations of occluded objects from context and commonsense knowledge, this capability remains a major challenge for vision-language models (VLMs). To address this gap, we introduce SceneFunRI, a benchmark for Reasoning the Invisible. Based on the SceneFun3D dataset, SceneFunRI formulates the task as a 2D spatial reasoning problem via a semi-automatic pipeline and comprises 855 instances. It requires models to infer the locations of invisible functional objects from task instructions and commonsense reasoning. The strongest baseline model (Gemini 3 Flash) only achieves an CAcc@75 of 15.20, an mIoU of 0.74, and a Dist of 28.65. We group our prompting analysis into three categories: Strong Instruction Prompting, Reasoning-based Prompting, and Spatial Process of Elimination (SPoE). These findings indicate that invisible-region reasoning remains an unstable capability in current VLMs, motivating future work on models that more tightly integrate task intent, commonsense priors, spatial grounding, and uncertainty-aware search.

视觉推理常识知识空间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。