arXiv:2410.11363cs.CV2024-10IJCV被引 2

通过人体-物体交互关系,提升物体使用可能性的识别能力。

Visual-Geometric Collaborative Guidance for Affordance Learning

  • 结合视觉与几何信息,从人-物交互中挖掘交互亲和性。
  • 在5.5万张图像数据集上实现更优的交互区域定位效果。
  • 适合研究物体功能理解与人机交互的学者参考。

感知图像中物体的潜在动作可能性(即使用性)并从人类示范中学习物体的交互功能,是一项因人-物交互多样性而极具挑战的任务。现有使用性学习方法多采用标签分配范式,假设功能区域与使用性标签存在唯一对应关系,在面对外观差异大的未见环境时表现不佳。本文提出利用交互亲和性进行使用性学习,即从人-物交互中提取交互亲和性,并将其迁移到非交互物体上。交互亲和性表征人体不同部位与目标物体局部区域间的接触关系,能提供人与物之间内在关联性的线索,从而降低对动作可能性感知的歧义性。为此,我们提出一种视觉-几何协同引导的使用性学习网络,联合利用视觉与几何线索挖掘人-物交互中的交互亲和性。此外,我们构建了一个接触驱动的使用性学习(CAL)数据集,包含超过55,047张图像的采集与标注。实验结果表明,该方法在客观指标和视觉质量上均优于代表性模型。

原文摘要 · Abstract (English)

Perceiving potential ``action possibilities'' (\ie, affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions. Prevailing affordance learning algorithms often adopt the label assignment paradigm and presume that there is a unique relationship between functional region and affordance label, yielding poor performance when adapting to unseen environments with large appearance variations. In this paper, we propose to leverage interactive affinity for affordance learning, \ie extracting interactive affinity from human-object interaction and transferring it to non-interactive objects. Interactive affinity, which represents the contacts between different parts of the human body and local regions of the target object, can provide inherent cues of interconnectivity between humans and objects, thereby reducing the ambiguity of the perceived action possibilities. To this end, we propose a visual-geometric collaborative guided affordance learning network that incorporates visual and geometric cues to excavate interactive affinity from human-object interactions jointly. Besides, a contact-driven affordance learning (CAL) dataset is constructed by collecting and labeling over 55,047 images. Experimental results demonstrate that our method outperforms the representative models regarding objective metrics and visual quality. Project: \href{https://github.com/lhc1224/VCR-Net}{github.com/lhc1224/VCR-Net}.

使用性学习人机交互几何引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。