用人类视角视频生成多样化机器人操作数据,提升跨体感泛化能力。
AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization

- 将人类交互分解为动作、视角、身体三因子,实现灵活重组
- 在RoboCasa和真实IRON机器人上提升操作性能,效果优于基线
- 无需配对数据,可生成带语言引导的空间目标选择数据
大规模收集丰富接触的机器人经验仍是实现可泛化操作的关键瓶颈。除数据量外,机器人学习还需跨体感、视角和场景的多样性体验。人类第一人称视频包含大量物理交互,但每段视频仅反映单一身体、相机轨迹与环境下的有限经验。我们提出AnyWorld,一种跨体感世界建模框架,可将单个真人交互扩展为多样化的机器人本体轨迹,无需成对的人-机示范。模型将交互分解为动作、相机和体感:动作控制运动结构,相机控制视角演化,目标体感上下文定义动作主体及其交互几何。该形式支持体感、视角与场景因子的独立重组,使单一模型生成多种机器人域经验,同时保留底层动力学与物体交互。模型采用大规模人类交互预训练+混合体感微调。实验表明,模型支持跨体感、视角与场景的可控重组;生成数据可提升RoboCasa GR1桌面试验及真实IRON人形机器人的操作表现。进一步验证发现,未配对的人类经验可重组为针对策略差距的机器人本体视频-动作对。受控的IRON干预纠正了虚假完成先验,建立了语言引导的空间目标选择;仅动作的反事实干预无法可靠学习后者,说明动作校准与视觉重构均不可或缺。
原文摘要 · Abstract (English)
Collecting contact-rich robot experiences at scale remains a major bottleneck for generalizable manipulation. Beyond data quantity, robot learning also requires diverse experiences across embodiments, viewpoints, and scenes. Human egocentric videos provide abundant physical interactions, but each video captures only a narrow slice of experience under a single body, camera trajectory, and environment. We propose AnyWorld, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations. Our model factorizes an interaction into action, camera, and embodiment: action controls capture the motion structure, camera controls specify viewpoint evolution, and the target embodiment context defines the acting body and its interaction geometry. This formulation enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain experiences while preserving the underlying dynamics and object interactions. We train the model with large-scale human interaction pretraining followed by mixed-embodiment fine-tuning. Experiments show that our model supports controllable recomposition across embodiments, viewpoints, and scenes, and we further demonstrate that the generated data can improve manipulation performance on the RoboCasa GR1 tabletop benchmark and a real IRON humanoid robot. Beyond aggregate gains, we test whether unpaired human experience can be recomposed into robot-native video-action pairs that target a policy gap. Controlled IRON interventions correct a spurious completion prior and establish language-grounded spatial target selection; an action-only counterfactual intervention fails to learn the latter reliably, showing that both action calibration and visual recomposition are necessary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。