从一段第三人称视频生成第一人称视角,让观众身临其境。
EgoX: Egocentric Video Generation from a Single Exocentric Video
- 用轻量LoRA适配大模型,融合第三人称与第一人称先验信息。
- 通过几何引导注意力机制,保持画面空间一致性与高画质。
- 仅需一个输入视频,即可生成逼真且泛化能力强的视角转换视频。
第一人称视角感知使人类能直接从自身视角理解世界。将第三人称视频转为第一人称视频,可带来沉浸式体验,但因相机姿态变化剧烈、视域重叠极少,仍极具挑战。该任务需在忠实保留可见内容的同时,合理合成未见区域,并保证几何一致性。为此,我们提出EgoX,一种从单个第三人称视频生成第一人称视频的新框架。EgoX通过轻量级LoRA适配大规模视频扩散模型的时空知识,引入统一条件策略,以通道与宽度拼接方式融合第三人称与第一人称先验。此外,设计几何引导自注意力机制,选择性关注空间相关区域,确保几何连贯性与高视觉保真度。方法在未见及真实场景视频上均表现良好,实现连贯且逼真的第一人称视频生成,具备强可扩展性与鲁棒性。
原文摘要 · Abstract (English)
Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibilities for immersive understanding but remains highly challenging due to extreme camera pose variations and minimal view overlap. This task requires faithfully preserving visible content while synthesizing unseen regions in a geometrically consistent manner. To achieve this, we present EgoX, a novel framework for generating egocentric videos from a single exocentric input. EgoX leverages the pretrained spatio temporal knowledge of large-scale video diffusion models through lightweight LoRA adaptation and introduces a unified conditioning strategy that combines exocentric and egocentric priors via width and channel wise concatenation. Additionally, a geometry-guided self-attention mechanism selectively attends to spatially relevant regions, ensuring geometric coherence and high visual fidelity. Our approach achieves coherent and realistic egocentric video generation while demonstrating strong scalability and robustness across unseen and in-the-wild videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。