用3D身体姿态控制第一人称视频生成,实现动作可预测。
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
- 通过3D全身姿态序列控制未来帧生成,结合相机运动与肢体动作
- 生成视频在时序上连贯、视觉真实,且与目标姿态高度一致
- 适合需要动作可控模拟的具身智能研究者
第一人称视频的细粒度动作控制是实现具身智能体模拟、预测与规划的关键需求。本文提出EgoControl,一种基于第一人称数据训练的姿态可控视频扩散模型。我们训练一个视频预测模型,使其根据显式的3D身体姿态序列生成未来帧。为实现精确运动控制,引入一种新型姿态表示,同时捕捉全局相机动态与关节式身体运动,并通过扩散过程中的专用控制机制进行融合。给定一段观察帧序列和目标姿态序列,EgoControl可生成与姿态控制对齐、时序连贯且视觉逼真的未来帧。实验表明,EgoControl能生成高质量、姿态一致的第一人称视频,为可控制的具身视频仿真与理解开辟新路径。
原文摘要 · Abstract (English)
Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion model trained on egocentric data. We train a video prediction model to condition future frame generation on explicit 3D body pose sequences. To achieve precise motion control, we introduce a novel pose representation that captures both global camera dynamics and articulated body movements, and integrate it through a dedicated control mechanism within the diffusion process. Given a short sequence of observed frames and a sequence of target poses, EgoControl generates temporally coherent and visually realistic future frames that align with the provided pose control. Experimental results demonstrate that EgoControl produces high-quality, pose-consistent egocentric videos, paving the way toward controllable embodied video simulation and understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。