发现视觉模型理解物体可用性的两大关键:形状结构与互动关系。
Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models
- 用探针分析视觉大模型,分离出几何与互动两类感知能力。
- 仅融合两种模型的特征,零样本即达到弱监督方法水平。
- 为视觉模型如何理解动作提供了可解释的机制框架。
视觉系统真正理解物体可用性意味着什么?我们认为这依赖于两种互补能力:几何感知——识别能引发交互的物体结构部件;以及互动感知——建模代理动作如何与这些部件作用。为验证此假设,我们对视觉基础模型(VFMs)进行了系统性探针分析。结果表明,DINO等模型天然编码了部件级几何结构,而Flux等生成模型则包含丰富的、受动词条件控制的空间注意力图,作为隐式的互动先验。关键的是,我们证明这两维并非仅相关,而是可组合的可用性构成要素。通过在无训练、零样本条件下简单融合DINO的几何原型与Flux的互动注意力图,即可实现媲美弱监督方法的可用性估计性能。该融合实验确认,几何与互动感知是视觉基础模型中可用性理解的根本构建块,为感知如何支撑行动提供了机制性解释。
原文摘要 · Abstract (English)
What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction, and interaction perception, which models how an agent's actions engage with those parts. To test this hypothesis, we conduct a systematic probing of Visual Foundation Models (VFMs). We find that models like DINO inherently encode part-level geometric structures, while generative models like Flux contain rich, verb-conditioned spatial attention maps that serve as implicit interaction priors. Crucially, we demonstrate that these two dimensions are not merely correlated but are composable elements of affordance. By simply fusing DINO's geometric prototypes with Flux's interaction maps in a training-free and zero-shot manner, we achieve affordance estimation competitive with weakly-supervised methods. This final fusion experiment confirms that geometric and interaction perception are the fundamental building blocks of affordance understanding in VFMs, providing a mechanistic account of how perception grounds action.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。