让第一视角视频与人体动作同步生成,真实还原佩戴者视角。
EgoTwin: Dreaming Body and View in First Person
- 用头部关节锚定人体动作,构建头中心运动表示
- 通过因果交互机制实现视频与动作的双向对齐
- 适合虚拟现实、智能穿戴等需要第一视角建模的场景
尽管外部视角视频合成已取得显著进展,第一人称视角视频生成仍研究不足,需同时建模第一视角内容与由穿戴者身体运动引发的相机运动轨迹。为弥合这一差距,我们提出联合生成第一人称视频与人体动作的新任务,面临两大挑战:1)视角对齐——生成视频中的相机轨迹必须与从人体动作推导出的头部轨迹精确一致;2)因果互动——合成的人体动作必须在时序上因果地匹配相邻视频帧间的视觉动态。为此,我们提出EgoTwin框架,基于扩散变换器架构,引入以头部关节为中心的运动表示,并设计受控制论启发的交互机制,在注意力运算中显式捕捉视频与动作间的因果关系。为全面评估,我们构建了一个大规模真实世界同步文本-视频-动作三元组数据集,并设计新指标评估视频与动作的一致性。大量实验验证了EgoTwin的有效性。
原文摘要 · Abstract (English)
While exocentric video synthesis has achieved great progress, egocentric video generation remains largely underexplored, which requires modeling first-person view content along with camera motion patterns induced by the wearer's body movements. To bridge this gap, we introduce a novel task of joint egocentric video and human motion generation, characterized by two key challenges: 1) Viewpoint Alignment: the camera trajectory in the generated video must accurately align with the head trajectory derived from human motion; 2) Causal Interplay: the synthesized human motion must causally align with the observed visual dynamics across adjacent video frames. To address these challenges, we propose EgoTwin, a joint video-motion generation framework built on the diffusion transformer architecture. Specifically, EgoTwin introduces a head-centric motion representation that anchors the human motion to the head joint and incorporates a cybernetics-inspired interaction mechanism that explicitly captures the causal interplay between video and motion within attention operations. For comprehensive evaluation, we curate a large-scale real-world dataset of synchronized text-video-motion triplets and design novel metrics to assess video-motion consistency. Extensive experiments demonstrate the effectiveness of the EgoTwin framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。