arXiv:2607.28243cs.CVcs.AI2026-07

用合成数据提升机器人操作泛化能力,生成逼真第一视角动作视频。

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

论文配图:EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
图 1 · 摘自论文原文
  • 通过动态3D场景锚点与动作导向位置编码,实现动作与环境的精准对齐。
  • 合成400条数据后,单臂任务成功率从77%提至84%,双臂任务从53%提至70%。
  • 适合需要扩展真实数据集的具身智能研究者,尤其关注第一视角动作生成。

第一视角视频为具身AI提供了丰富的操作体验,但跨场景、物体、动作和主体收集多样化数据成本高昂。本文提出 extit{EgoGenesis},一个第一视角世界-动作模拟器,可生成可控的高质量操作视频,以扩充稀缺的真实训练数据。该模型基于预训练视频生成先验,引入两种几何感知的条件机制:在线锚定投影记忆(OAPM)在自回归生成中保留首帧3D场景锚点,并周期性刷新近期状态;动作-3D旋转变换位置编码(A3D-RoPE)使用相机感知的3D旋转坐标编码末端执行器运动,将动作几何信息注入骨架到视频的交叉注意力中,实现精确控制。二者协同提升了长序列第一视角生成的视觉保真度、几何稳定性与动作一致性。此外,将400条真实轨迹与400条 extit{EgoGenesis}生成轨迹结合,在单臂任务上使真实机器人分布外成功率从77%提升至84%,双臂任务从53%提升至70%,证明合成数据显著增强下游世界动作建模(WAM)的泛化能力。

原文摘要 · Abstract (English)

Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory (OAPM) preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding (A3D-RoPE) encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control. Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Moreover, augmenting 400 real trajectories with 400 \method-generated trajectories improves out-of-distribution real-robot success from 77\% to 84\% on single-arm tasks and from 53\% to 70\% on dual-arm tasks, demonstrating that the synthesized data substantially improve downstream WAM generalization.

第一视角动作生成具身智能数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。