arXiv:2409.06210cs.CV2024-09ECCV被引 16

用外视角图像无监督学习物体交互特征,无需成对数据

INTRA: Interaction Relationship-aware Weakly Supervised Affordance Grounding

  • 仅用外视角图像做对比学习,提取交互特异性特征
  • 在多个数据集上超越现有方法,合成图也能泛化
  • 支持任意文本条件生成交互图,适合新物体新动作

affordance 指物体固有的潜在交互能力。感知 affordance 可使智能体高效导航并交互新环境。弱监督 affordance grounding 在无需昂贵像素级标注的情况下教会代理理解 affordance,但依赖外视角图像。尽管近期进展取得良好效果,仍面临需成对的外视角与内视角图像数据集、以及单个物体多种 affordance 难以定位等挑战。为此,我们提出 INTRA:Interaction Relationship-aware Weakly Supervised Affordance Grounding。不同于以往方法,INTRA 将问题重定义为仅使用外视角图像进行表示学习,通过对比学习识别交互的唯一特征,彻底消除对配对数据集的需求。此外,我们利用视觉-语言模型嵌入,实现任意文本条件下的 affordance grounding,设计文本条件交互图生成以反映交互关系,并通过文本同义词增强提升鲁棒性。在 AGD20K、IIT-AFF、CAD、UMD 等多样数据集上,我们的方法优于先前方法。实验还表明,该方法在合成图像/插图上具有显著领域可扩展性,能对新交互和新物体完成 affordance grounding。

原文摘要 · Abstract (English)

Affordance denotes the potential interactions inherent in objects. The perception of affordance can enable intelligent agents to navigate and interact with new environments efficiently. Weakly supervised affordance grounding teaches agents the concept of affordance without costly pixel-level annotations, but with exocentric images. Although recent advances in weakly supervised affordance grounding yielded promising results, there remain challenges including the requirement for paired exocentric and egocentric image dataset, and the complexity in grounding diverse affordances for a single object. To address them, we propose INTeraction Relationship-aware weakly supervised Affordance grounding (INTRA). Unlike prior arts, INTRA recasts this problem as representation learning to identify unique features of interactions through contrastive learning with exocentric images only, eliminating the need for paired datasets. Moreover, we leverage vision-language model embeddings for performing affordance grounding flexibly with any text, designing text-conditioned affordance map generation to reflect interaction relationship for contrastive learning and enhancing robustness with our text synonym augmentation. Our method outperformed prior arts on diverse datasets such as AGD20K, IIT-AFF, CAD and UMD. Additionally, experimental results demonstrate that our method has remarkable domain scalability for synthesized images / illustrations and is capable of performing affordance grounding for novel interactions and objects.

弱监督交互感知视觉语言泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。