用多源第一视角数据训练机器人,实现更灵活的抓取与操作能力。
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
- 融合人类与机器人的第一视角数据,统一动作空间
- 在6个真实任务中平均成功率最高,泛化能力强
- 适合希望提升机器人灵巧操作能力的研究者
构建能跨任务感知、推理与执行的通用机器人仍是一大挑战,尤其在灵巧操作方面。主要瓶颈在于缺乏大规模、带动作标注的灵巧技能数据,因遥操作成本高且困难。人类数据具有规模大、行为多样等优势,可为机器人学习提供丰富先验。尽管已有研究尝试利用人类示范,但受限于场景有限及人机间视觉差异。为此,我们提出METIS,一个基于多源第一视角数据预训练的视觉-语言-动作(VLA)模型。首先构建EgoAtlas,整合来自多个来源的人类与机器人数据,统一至一致的动作空间。进一步提取运动感知动态,一种紧凑且离散化的运动表示,为VLA训练提供高效表达的监督信号。在此基础上,METIS将推理与执行融合于统一框架,有效部署于下游灵巧操作任务。实验表明,该方法在六个真实世界任务中取得最高平均成功率,且对分布外场景表现出优异的泛化与鲁棒性。这些结果表明METIS是迈向灵巧操作通用模型的重要一步。
原文摘要 · Abstract (English)
Building a generalist robot that can perceive, reason, and act across diverse tasks remains an open challenge, especially for dexterous manipulation. A major bottleneck lies in the scarcity of large-scale, action-annotated data for dexterous skills, as teleoperation is difficult and costly. Human data, with its vast scale and diverse manipulation behaviors, provides rich priors for learning robotic actions. While prior works have explored leveraging human demonstrations, they are often constrained by limited scenarios and a large visual gap between human and robots. To eliminate these limitations, we propose METIS, a vision-language-action (VLA) model for dexterous manipulation pretrained on multi-source egocentric datasets. We first construct EgoAtlas, which integrates large-scale human and robotic data from multiple sources, all unified under a consistent action space. We further extract motion-aware dynamics, a compact and discretized motion representation, which provides efficient and expressive supervision for VLA training. Built upon them, METIS integrates reasoning and acting into a unified framework, enabling effective deployment to downstream dexterous manipulation tasks. Our method demonstrates exceptional dexterous manipulation capabilities, achieving highest average success rate in six real-world tasks. Experimental results also highlight the superior generalization and robustness to out-of-distribution scenarios. These findings emphasize METIS as a promising step toward a generalist model for dexterous manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。