让第一视角视频生成更真实可控,支持自身体态与他人动作的精确控制。
E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control

- 用3D环境记忆+人体姿态渲染,分离场景结构与动态变化
- 在Nymeria数据集上显著提升画面质量与动作一致性
- 适合需要精准第一视角生成的智能体、虚拟现实应用
可控且物理合理的第一视角视频生成对具身智能体理解自身与他人行为如何改变世界至关重要。相比通用视频合成,第一视角生成更具挑战:相机与主体紧密耦合,导致视角快速变化和频繁自遮挡;动作细微、复杂且常部分可见;人物与场景状态需一致响应控制指令。我们提出E³C,一种可控的第一视角视频扩散框架,通过结构化、紧凑的条件设计,解耦持久场景结构与人驱动动态。从上下文帧中,E³C构建基于半密集点云的3D记忆,并为每个点附加来自视频VAE特征的外观描述符。将该记忆渲染至目标视角,生成与目标帧对齐的条件。人体动态独立建模:场景中观察到的人体由骨架渲染控制(外视角人体控制),而摄像头佩戴者则由其3D关节和6自由度手腕运动控制(内视角人体控制)。为在佩戴者肢体被遮挡时仍保持内视角控制,引入一个持续的跨注意力编码器。在Nymeria数据集上的实验表明,E³C在视觉保真度、相机运动准确性、物体一致性以及内外视角人体控制方面优于强基线模型,同时支持直观的场景编辑。
原文摘要 · Abstract (English)
Controllable and physically grounded egocentric video generation is essential for embodied agents to reason about how their own and others' actions manifest and change the world. Compared to generic video synthesis, egocentric generation is especially challenging: the camera is tightly coupled to the actor, leading to rapid viewpoint changes and frequent self-occlusions; the underlying actions are subtle, articulated, and often only partially visible; and both the people and the scene state must evolve consistently with the specified controls. We present E$^3$C, a controllable video diffusion framework for egocentric generation that builds structured and compact conditions disentangling persistent scene structure from human-driven dynamics. From context frames, E$^3$C constructs a semi-dense point cloud-based 3D memory and augments each point with appearance descriptors from video-VAE features. Rendering this memory into target viewpoints produces conditioning aligned with the target frames. Human dynamics are modeled separately. The observed people in the scene are controlled by skeleton renderings (exo human control), while the camera wearer is specified by their 3D body joints and 6DoF wrist motion (ego human control). To preserve ego human control when the wearer's body parts are invisible, we introduce an ego motion encoder that produces persistent cross-attention tokens. Experiments on Nymeria show that E$^3$C improves visual fidelity, camera-motion accuracy, object consistency, and ego & exo human control over strong baselines, while also enabling intuitive scene editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。