通过反事实补全揭示图像中物体间的依赖关系
Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting
- 利用物体间不对称关系与大模型补全生成反事实图
- 无需训练即可在真实图像上有效识别可移除物体
- 适合研究场景理解与视觉推理的学者
本文提出一种新型场景理解任务——视觉积木(Visual Jenga)。受游戏积木启发,该任务通过逐步移除图像中的物体直至仅剩背景,以揭示场景元素间的内在依赖关系。正如积木玩家需理解结构依赖以维持塔体稳定,本任务通过系统性探索哪些物体可被移除而仍保持场景在物理和几何层面的连贯性来实现。为启动该任务,我们提出一种简单、数据驱动且无需训练的方法,在多种真实图像上表现惊人有效。其核心原理是利用场景内物体间配对关系的不对称性,并借助大型补全模型生成一组反事实图像以量化这种不对称性。
原文摘要 · Abstract (English)
This paper proposes a novel scene understanding task called Visual Jenga. Drawing inspiration from the game Jenga, the proposed task involves progressively removing objects from a single image until only the background remains. Just as Jenga players must understand structural dependencies to maintain tower stability, our task reveals the intrinsic relationships between scene elements by systematically exploring which objects can be removed while preserving scene coherence in both physical and geometric sense. As a starting point for tackling the Visual Jenga task, we propose a simple, data-driven, training-free approach that is surprisingly effective on a range of real-world images. The principle behind our approach is to utilize the asymmetry in the pairwise relationships between objects within a scene and employ a large inpainting model to generate a set of counterfactuals to quantify the asymmetry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。