通过结构化语义锚定,提升极端视角下视频生成质量。
Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation

- 双分支扩散模型:结合几何重建与物体级语义上下文
- 在EgoExo4D上显著超越现有方法,全指标领先
- 自动生成合成数据,解决相机-重建错位问题
从单个外视角视频生成第一人称视频是AR/VR与物理AI中的重要课题。与传统新视角合成相比,外到内生成因视角剧变和大量不可见区域导致几何条件失效,难度更高。我们提出Grounded-Exo2Ego,一种在架构与数据层面均具原则性的框架。架构上采用双分支视频扩散模型:一个几何锚定分支基于3D重建渲染进行条件生成,一个新颖的语义锚定分支则利用物体级上下文合成挑战性区域,突破现有几何依赖范式。此外,我们发现相机-重建错位严重损害训练效果,因而引入相机重定位算法,显著提升各项指标。我们还构建全自动合成数据引擎,生成并渲染程序化环境中的绑定3D角色。在挑战性EgoExo4D数据集上的评估显示,本方法在所有指标上大幅优于近期最先进方法。详细消融实验验证了架构与数据层面各贡献的有效性。
原文摘要 · Abstract (English)
Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego generation is a significantly harder task because the standard geometric conditioning becomes highly unreliable under extreme view changes and large unobservable regions. We present Grounded-Exo2Ego, a principled framework that addresses these challenges at both the architectural and data levels. Architecturally, Grounded-Exo2Ego is a dual-branch video diffusion model that couples a geometric anchoring branch, which conditions the generation on the rendering of a 3D reconstruction, with a novel semantic grounding branch, which goes beyond the prevailing geometry-based approach and improves quality by synthesizing challenging regions based on object-level context. Additionally, we found that the overlooked issue of camera-reconstruction misalignment severely undermines exo-to-ego learning. We thus introduce a camera re-localization algorithm that resolves this issue and substantially improves quality across all metrics. We further develop a fully automated synthetic data engine that generates and renders rigged 3D characters in procedurally generated environments. Evaluation on the challenging EgoExo4D dataset shows that our method outperforms recent state-of-the-art approaches by large margins across all metrics. Detailed ablations validate improvements from each of our contributions at both the data and architectural level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。