arXiv:2511.15704cs.ROcs.AI2025-11被引 25

用真实场景与任务数据训练出能听懂指令的机械臂,提升泛化能力。

In-N-On: Scaling Egocentric Manipulation with in-the-wild and on-task Data

  • 区分真实场景与任务数据,构建大规模混合数据集PHSD
  • 模型仅用人类数据即可实现语言指令跟随和少样本学习
  • 适合想提升机器人操作泛化能力的研究者与工程师

第一人称视角视频是学习操控策略的宝贵且可扩展的数据来源。然而由于数据异质性显著,现有方法大多仅用人类数据进行简单预训练,未能发挥其全部潜力。本文首次提出一套可扩展的数据收集与使用方案,将人类数据分为真实场景(in-the-wild)和任务对齐(on-task)两类,并系统分析其使用方式。我们构建了PHSD数据集,包含超过1000小时多样化的第一人称真实场景数据和超过20小时直接对齐目标操控任务的标注数据。基于此,我们训练了一个大型第一人称语言条件流匹配策略模型Human0。通过领域自适应技术,Human0有效缩小了人类与类人机器人之间的差距。实验表明,仅通过扩大人类数据规模,Human0实现了语言指令跟随、少样本学习及使用任务数据后鲁棒性的提升等新特性。

原文摘要 · Abstract (English)

Egocentric videos are a valuable and scalable data source to learn manipulation policies. However, due to significant data heterogeneity, most existing approaches utilize human data for simple pre-training, which does not unlock its full potential. This paper first provides a scalable recipe for collecting and using egocentric data by categorizing human data into two categories: in-the-wild and on-task alongside with systematic analysis on how to use the data. We first curate a dataset, PHSD, which contains over 1,000 hours of diverse in-the-wild egocentric data and over 20 hours of on-task data directly aligned to the target manipulation tasks. This enables learning a large egocentric language-conditioned flow matching policy, Human0. With domain adaptation techniques, Human0 minimizes the gap between humans and humanoids. Empirically, we show Human0 achieves several novel properties from scaling human data, including language following of instructions from only human data, few-shot learning, and improved robustness using on-task data. Project website: https://xiongyicai.github.io/In-N-On/

机器人操控第一人称视觉语言指令少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。