arXiv:2503.09320cs.CVcs.LG2025-03ICCV被引 15

从人类视频中学习双手操作的精准物体可用区域,让机器人更懂怎么用手做事。

2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos

  • 基于人类动作视频提取双手协同操作的物体可用区域
  • 构建2HANDS数据集,包含精细分割与任务标签的双臂操作标注
  • 模型可生成机器人可用的操作区域,适用于抓取、翻转等实际任务

在与物体交互时,人类能准确判断物体哪些区域适合执行特定动作,即物体的可用性区域。他们还能根据任务需求和单手或双手使用情况,区分物体区域的细微差异。然而,当前基于视觉的可用性预测方法常简化为对象部件分割。本文提出一种从人类活动视频数据集中提取可用性数据的框架。我们构建的2HANDS数据集包含精确的对象可用性区域分割结果和动作类别标签,明确描述了所执行的动作。该数据还涵盖双臂协同操作,即两隻手协调作用于一个或多个物体。我们提出基于视觉语言模型的2HandedAfforder模型,在该数据集上训练并展示了在多种动作中优于基线的可用性区域分割性能。最后,我们在机器人操作场景中验证了预测的可用性区域具有可操作性,能够指导机器人完成任务。项目网站:https://sites.google.com/view/2handedafforder

原文摘要 · Abstract (English)

When interacting with objects, humans effectively reason about which regions of objects are viable for an intended action, i.e., the affordance regions of the object. They can also account for subtle differences in object regions based on the task to be performed and whether one or two hands need to be used. However, current vision-based affordance prediction methods often reduce the problem to naive object part segmentation. In this work, we propose a framework for extracting affordance data from human activity video datasets. Our extracted 2HANDS dataset contains precise object affordance region segmentations and affordance class-labels as narrations of the activity performed. The data also accounts for bimanual actions, i.e., two hands co-ordinating and interacting with one or more objects. We present a VLM-based affordance prediction model, 2HandedAfforder, trained on the dataset and demonstrate superior performance over baselines in affordance region segmentation for various activities. Finally, we show that our predicted affordance regions are actionable, i.e., can be used by an agent performing a task, through demonstration in robotic manipulation scenarios. Project-website: https://sites.google.com/view/2handedafforder

机器人操作双臂协作视觉理解动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。