arXiv:2607.08436cs.ROcs.AI2026-07被引 5

用人体视角数据训练机器人,让模型预测场景变化而非只模仿动作。

EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data

论文配图:EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
图 1 · 摘自论文原文
  • 通过预测场景演变来训练策略,分离可迁移与不可迁移因素。
  • 相比传统模仿学习,使用DINO和3D运动流提升性能达20%-4倍。
  • 适合研究人机共训、视觉表征迁移与真实世界机器人操作的学者。

第一人称人类数据为机器人操作提供了可扩展的监督信号。然而,行为克隆将可迁移的内容(如物体、场景、任务语义)与不可迁移的因素(如人体结构、头部运动、行为风格)纠缠在一起。本文探讨世界动作模型(WAMs)是否能提供更优的训练信号,即要求策略不仅预测动作,还需预测场景演化。核心问题是:何种世界表征最利于人类到机器人的迁移?我们假设有效的世界目标应抽象外观、捕捉与代理无关的物理效应,并分离相机运动与环境变化。为此,提出EgoWAM,一个控制变量的人机联合训练框架,固定策略主干、动作头和数据混合比例,仅改变世界预测目标(像素、DINO、3D运动流)。在三个真实世界双臂任务中,WAM联合训练比行为克隆更有效利用野外第一人称数据。基于像素的预测迁移能力弱,而DINO和3D流带来显著提升:DINO使分布外的物体与场景泛化能力提升至4倍,3D流在域内表现提升20%-30%。

原文摘要 · Abstract (English)

Egocentric human data offers scalable supervision for robot manipulation. However, behavior cloning entangles transferable content like objects, scenes, and task semantics, with non-transferable factors like human morphology, head motion, and behavioral style. We study whether World Action Models (WAMs) provide a better training signal by requiring policies to predict not only actions, but also how the scene evolves. The central question is what world representation best enables human-to-robot transfer. We hypothesize that an effective world target should abstract appearance, capture agent-invariant physical effects, and separate camera motion from environment change. We introduce EgoWAM, a controlled human-robot co-training framework that fixes the policy backbone, action head, and data mixture while varying only the world prediction target, comparing Pixel, DINO, and 3D motion flow. Across three real-world bimanual tasks, WAM co-training scales more effectively with in-the-wild egocentric human data than behavior cloning. Pixel-based prediction transfers weakly, while DINO and 3D flow yield substantial gains: DINO improves out-of-distribution object and scene generalization by up to 4x, and 3D flow improves in-domain performance by 20-30%. More details: https://gatech-rl2.github.io/egowam.github.io

机器人操作世界模型第一人称数据跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。