arXiv:2510.17384cs.CV2025-10ICCV被引 5

让视觉模型在看别人用东西时学会自己用,还能反过来优化理解能力。

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

  • 双向知识循环:从旁观视频学动作区域,再反向提升理解力
  • 在物体被人体完全遮挡的复杂场景中仍能准确定位可操作区域
  • 适合做智能机器人交互、人机协作等需要动态理解的场景

人类仅通过观察他人使用新物品,就能完成此前未经历的交互。弱监督的可操作性定位模仿这一过程,在仅有图像级标注的非视角交互图像上训练,以定位自我视角图像中支持动作的物体区域。然而,以往方法仅单向从非视角图像迁移知识到自我视角图像,限制了其在复杂交互场景中的应用。本文提出闭环迁移框架 LoopTrans,不仅将知识从非视角传至自我视角,还实现反向传递以增强非视角知识提取。LoopTrans引入统一跨模态定位与去噪知识蒸馏等创新机制,弥合以物体为中心的自我视角与以交互为中心的非视角图像之间的域差距,提升知识迁移效果。实验表明,LoopTrans在图像与视频基准测试中各项指标均持续提升,甚至在物体交互区域被人体完全遮挡的挑战性场景中表现优异。

原文摘要 · Abstract (English)

Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions that enable actions on egocentric images, using exocentric interaction images with image-level annotations. However, extracting affordance knowledge solely from exocentric images and transferring it one-way to egocentric images limits the applicability of previous works in complex interaction scenarios. Instead, this study introduces LoopTrans, a novel closed-loop framework that not only transfers knowledge from exocentric to egocentric but also transfers back to enhance exocentric knowledge extraction. Within LoopTrans, several innovative mechanisms are introduced, including unified cross-modal localization and denoising knowledge distillation, to bridge domain gaps between object-centered egocentric and interaction-centered exocentric images while enhancing knowledge transfer. Experiments show that LoopTrans achieves consistent improvements across all metrics on image and video benchmarks, even handling challenging scenarios where object interaction regions are fully occluded by the human body.

动作理解闭环学习弱监督视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。