arXiv:2605.15477cs.CV2026-05

用外视角视频生成内视角动作表示,提升机器人规划能力

EgoExo-WM: Unlocking Exo Video for Ego World Models

论文配图:EgoExo-WM: Unlocking Exo Video for Ego World Models
图 1 · 摘自论文原文
  • 从外视角视频提取身体姿态作为动作表征,转换为内视角视频
  • 用转换数据训练世界模型,预测与规划性能显著提升
  • 适合研究机器人视觉、动作建模及增强现实的开发者

内视角世界模型有望推动智能体的预测与规划,但受限于内视角训练数据稀缺及其对人类动作的观测不全。相比之下,外视角视频丰富且能清晰捕捉人体姿态,却缺乏与智能体动作空间的直接对应关系。本文提出一种方法,通过提取外视角视频中的结构化身体姿态作为动作表征,并借助人体运动学先验将外视角视频转换为内视角视频,从而实现真实场景外视角数据在内视角世界模型训练中的可用性。实验表明,使用该转换数据训练全身动作条件下的内视角世界模型,显著提升了预测精度与下游规划性能,可推断达成视觉目标状态所需的身体姿态序列。本方法为利用任意真实场景外视角视频构建强大内视角世界模型开辟了新路径,有助于推进机器人规划与增强现实引导等应用。

原文摘要 · Abstract (English)

Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric training data and its inherent partial observability of humans' physical actions. In contrast, exocentric video is abundant and reveals body poses well, but lacks direct alignment with an agent's action space -- and is not egocentric. We propose a method to bridge this gap by extracting structured body pose from exocentric video as a representation of action and transforming the exocentric video to egocentric video, informed by a human kinematics prior. This process unlocks the integration of in-the-wild exocentric data for egocentric world model training. We show that training whole-body action-conditioned egocentric world models with our converted data significantly improves both prediction quality and downstream planning performance, where we infer the sequence of body poses needed to achieve a visual goal state. Our approach paves the way to enlist arbitrary in-the-wild videos for building powerful egocentric world models, furthering applications in robot planning and augmented-reality guidance.

世界模型动作生成视频转换机器人规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。