让视频生成模型从第三人称视角生成第一人称视频
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
- 通过视角对齐、多视角融合和姿态注入三模块,实现跨视角视频生成
- 在ExoEgo4D数据集上显著提升第一人称视频生成质量
- 无需重新训练即可将现有模型用于跨视角生成,适合视觉生成研究者
以WAN 2.2为代表的基础视频生成模型具备强大的文本与图像条件生成能力,但仅限于同视角生成。本文提出Exo2EgoSyn,将WAN 2.2适配为第三人称到第一人称(Exo2Ego)跨视角视频生成框架。该框架包含三个关键模块:Ego-Exo视图对齐(EgoExo-Align)在潜在空间对齐第三人称与第一人称首帧表征,引导生成空间从第三人称转向第一人称;多视角第三人称视频条件化(MultiExoCon)将多视角第三人称视频融合为统一条件信号,扩展了WAN2.2的单图或文本条件限制;姿态感知潜在注入(PoseInj)将相对第三人称到第一人称的相机位姿信息注入潜在状态,指导跨视角的几何感知生成。三模块协同实现仅凭第三人称观测生成高质量第一人称视频,且无需从头训练。在ExoEgo4D数据集上的实验表明,Exo2EgoSyn显著提升第一人称生成效果,为基于基础模型的可扩展跨视角视频生成开辟新路径。源代码与模型将公开发布。
原文摘要 · Abstract (English)
Foundation video generation models such as WAN 2.2 exhibit strong text- and image-conditioned synthesis abilities but remain constrained to the same-view generation setting. In this work, we introduce Exo2EgoSyn, an adaptation of WAN 2.2 that unlocks Exocentric-to-Egocentric(Exo2Ego) cross-view video synthesis. Our framework consists of three key modules. Ego-Exo View Alignment(EgoExo-Align) enforces latent-space alignment between exocentric and egocentric first-frame representations, reorienting the generative space from the given exo view toward the ego view. Multi-view Exocentric Video Conditioning (MultiExoCon) aggregates multi-view exocentric videos into a unified conditioning signal, extending WAN2.2 beyond its vanilla single-image or text conditioning. Furthermore, Pose-Aware Latent Injection (PoseInj) injects relative exo-to-ego camera pose information into the latent state, guiding geometry-aware synthesis across viewpoints. Together, these modules enable high-fidelity ego view video generation from third-person observations without retraining from scratch. Experiments on ExoEgo4D validate that Exo2EgoSyn significantly improves Ego2Exo synthesis, paving the way for scalable cross-view video generation with foundation models. Source code and models will be released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。