arXiv:2509.17074cs.CVcs.AI2025-09

用信息约束提升图文对齐,让机器人更好理解物体使用方式。

Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models

  • 通过最大化图文特征互信息,实现精准的文本引导定位。
  • 在AGD20K数据集上,单样本学习性能达新基准。
  • 适合做视觉可操作性理解与少样本视觉推理的研究者。

视觉可操作性学习对机器人理解并有效交互物理世界至关重要。近期方法利用预训练的视觉语言基础模型,在有限数据下学习可操作属性,开创了新范式。但这些方法忽视了图像与文本描述在特征层面的对齐重要性,导致识别效果不佳。本文提出一种基于信息约束的文本引导可操作性学习框架。设计可操作性互信息约束,通过最大化输入图像中可操作区域特征与对应文本提示间的互信息,同步优化文本提示与任务相关的视觉特征。同时提出对象级信息约束,最大化给定对象的视觉特征与其所属类别文本特征之间的互信息,以获取高质量对象表示,为可操作区域识别提供可靠语义先验。在AGD20K数据集上的实验表明,该方法优于现有方法,实现了单样本可操作性学习的新最佳性能。

原文摘要 · Abstract (English)

Visual affordance learning is crucial for robots to understand and interact effectively with the physical world. Recent advances in this field attempt to leverage pre-trained knowledge of vision-language foundation models to learn affordance properties with limited training data, providing a novel paradigm for visual affordance learning. However, these methods overlook the significance of maintaining feature alignment between visual images and language descriptions for identifying affordance areas with textual guidance, and thus may lead to suboptimal results. In this paper, we present an informative framework for text-guided affordance learning, which involves information-based constraints to achieve text-image alignment at feature level. Specifically, we design an affordance mutual information constraint that helps learn appropriate textual prompts and task-oriented visual features simultaneously by maximizing the mutual information between the features of the affordance areas in the input images and the corresponding textual prompts. In addition, we propose an object-level information constraint that maximizes the mutual information between the visual features of a given object and the text features of the category it belongs to. This enables the model to capture high-quality representations for the object, providing more reliable semantic priors for identifying affordance regions. Experimental results on the AGD20K dataset show that the proposed method outperforms existing approaches and achieves the new state-of-the-art in one-shot affordance learning.

视觉可操作性图文对齐少样本学习基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。