arXiv:2505.05074cs.CVcs.RO2025-05综述被引 2

统一视觉可操作性预测定义,提升方法可比性与复现性

Visual Affordance Prediction: Survey and Reproducibility

  • 提出统一的视觉可操作性预测框架,涵盖物体与主体交互全信息
  • 揭示现有方法与数据集在定义上的不一致导致比较失真
  • 推出可操作性清单,推动研究透明化和公平性

可操作性指代理通过摄像头观察到的对物体可能执行的动作。视觉可操作性预测在抓取检测、可操作性分类、可操作性分割和手部姿态估计等任务中存在不同表述,导致定义不一致,阻碍方法间的公平比较。本文通过综合考虑目标物体的完整信息及代理与物体的交互过程,提出一种统一的视觉可操作性预测形式化方法。该统一框架使我们能够系统性地回顾分散的视觉可操作性研究,揭示方法与数据集的优势与局限。同时,论文讨论了可复现性问题,如方法实现与实验设置细节缺失,使得基准测试不公平且不可靠。为促进透明度,我们引入可操作性清单(Affordance Sheet),详细记录方法、数据集与验证流程,支持未来研究的可复现性与公平性。

原文摘要 · Abstract (English)

Affordances are the potential actions an agent can perform on an object, as observed by a camera. Visual affordance prediction is formulated differently for tasks such as grasping detection, affordance classification, affordance segmentation, and hand pose estimation. This diversity in formulations leads to inconsistent definitions that prevent fair comparisons between methods. In this paper, we propose a unified formulation of visual affordance prediction by accounting for the complete information on the objects of interest and the interaction of the agent with the objects to accomplish a task. This unified formulation allows us to comprehensively and systematically review disparate visual affordance works, highlighting strengths and limitations of both methods and datasets. We also discuss reproducibility issues, such as the unavailability of methods implementation and experimental setups details, making benchmarks for visual affordance prediction unfair and unreliable. To favour transparency, we introduce the Affordance Sheet, a document that details the solution, datasets, and validation of a method, supporting future reproducibility and fairness in the community.

可操作性可复现性统一框架综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。