arXiv:2505.11865cs.ROcs.CV2025-05被引 18

用50万张带可操作性提示的图像训练机器人,让其从人类动作中学会通用抓取技能。

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation

  • 基于50万张图像构建大规模可操作性数据集HOVA-500K,覆盖1726类物体和675种动作。
  • 提出GLOVER++框架,在新任务上实现领先性能,跨场景、跨模态泛化能力强。
  • 适合研究人机交互、具身智能与机器人抓取的学者与工程师使用。

从人类示范视频学习操作技能为实现可泛化且可解释的机器人智能提供了可行路径,尤其通过可操作性(affordance)的视角。然而,知识迁移仍面临两大挑战:一是缺乏大规模且精确标注可操作性的数据集;二是对多样化操作情境下可操作性的探索不足。为此,我们构建了HOVA-500K,一个包含50万张图像、覆盖1726类物体和675种动作的大规模可操作性标注数据集,并发布了标准化的多模态可操作性推理基准。在此基础上,我们提出了GLOVER++——一种全局到局部的可操作性训练框架,能有效将人类示范中的可操作性知识迁移到下游开放词汇推理任务中。GLOVER++在HOVA-500K基准上取得当前最优结果,并在多种下游机器人操作任务中展现出强泛化能力。通过显式建模可操作性,GLOVER++实现了跨场景、跨模态、跨任务的稳健迁移。我们希望HOVA-500K与GLOVER++框架能成为连接人类示范与机器人操作能力的重要资源。

原文摘要 · Abstract (English)

Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence-particularly through the lens of actionable affordances. However, transferring such knowledge remains challenging due to: 1) a lack of large-scale datasets with precise affordance annotations, and 2) insufficient exploration of affordances in diverse manipulation contexts. To address these gaps, we introduce HOVA-500K, a large-scale, affordance-annotated dataset comprising 500,000 images across 1,726 object categories and 675 actions. We also release a standardized benchmarking suite for multi-modal affordance reasoning. Built upon HOVA-500K, we present GLOVER++, a global-to-local affordance training framework that effectively transfers actionable affordance knowledge from human demonstrations to downstream open-vocabulary reasoning tasks. GLOVER++ achieves state-of-the-art results on the HOVA-500K benchmark and demonstrates strong generalization across diverse downstream robotic manipulation tasks. By explicitly modeling actionable affordances, GLOVER++ facilitates robust transfer across scenes, modalities, and tasks. We hope that HOVA-500K and the GLOVER++ framework will serve as valuable resources for bridging the gap between human demonstrations and robotic manipulation capabilities.

机器人操作可操作性视觉理解人类示范

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。