arXiv:2608.05215cs.ROcs.CV2026-08中稿 · IEEE/RSJ IROS 2026

用人类视频学物体可操作性,让机器人直接套用。

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

论文配图:VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances
图 1 · 摘自论文原文
  • 从第一视角视频提取视觉、抓取、运动三类可操作性信息
  • 构建含560万视觉、1160万抓取/轨迹的大型数据集
  • 模型能零样本生成机器人可执行的动作指令

从人类视频中学习操作技能是实现机器人可扩展学习的有前景方向。然而,人与机器人的身体差异使该任务面临挑战。一种可行方案是学习与具体身体无关的、以物体为中心的可操作性。本文提出框架,利用先进3D结构光与手部网格重建技术,从第一视角人类视频中提取视觉、抓取和轨迹等可操作性信息,明确指示交互位置、抓取方式和运动路径。我们构建了EgoAffordance数据集,包含20.4万段视频,涵盖560万条视觉可操作性、1160万条抓取与轨迹可操作性。在此基础上,提出VLAff——基于大规模视觉-语言模型的统一基础模型,学习各类可操作性间的跨模态关联。给定视觉观察与指令,VLAff生成视觉可操作性热图、抓取姿态和轨迹,并结合3D场景信息转化为可直接执行的动作。大量实验证明,VLAff在视觉可操作性预测上达到当前最优性能,且可有效应用于真实机器人任务,如零样本操作和可操作性引导的学习。

原文摘要 · Abstract (English)

Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One promising solution is to learn object-centric actionable affordances that are embodiment-agnostic. In this work, we propose a framework that leverages egocentric human videos with state-of-the-art 3D Structure-from-Motion and hand mesh reconstruction to extract actionable affordances such as visual, grasp, and trajectory affordances that explicitly encode where to interact, how to grasp, and how to move. We construct EgoAffordance, a large-scale dataset comprising 204K episodes with 5.6M visual affordances and 11.6M grasp and trajectory affordances. Building on this, we introduce VLAff, a large vision-language model-based unified foundation model that learns cross-modal correlations across all actionable affordances. Given a visual observation and instruction, VLAff generates visual affordance heatmaps, grasp poses, and trajectories, which are then converted into directly executable actions by utilizing 3D scene information. Through extensive experiments, we demonstrate that VLAff not only achieves state-of-the-art performance on visual affordance prediction, but can also be effectively applied to real robot applications such as zero-shot manipulation and affordance-guided robot learning.

机器人学习可操作性视觉-语言模型动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。