arXiv:2506.05554cs.CV2025-06被引 33

用深度闭合网格生成极端视角4D视频,解决几何失真问题

EX-4D: EXtreme Viewpoint 4D Video Synthesis via Depth Watertight Mesh

  • 用深度闭合网格建模可见与被遮区域,保持极端视角几何一致
  • 仅用单目视频通过模拟遮挡生成训练数据,无需多视角配对
  • 结合轻量LoRA适配器,实现物理一致且时间连贯的高质量视频

从单目输入生成可控制相机视角的高质量视频是一项挑战,尤其在极端视角下。现有方法常因几何不一致和边界遮挡伪影导致视觉质量下降。本文提出EX-4D,通过深度闭合网格表示作为鲁棒几何先验,显式建模可见与被遮区域,确保极端相机姿态下的几何一致性。为克服缺乏成对多视角数据的问题,提出一种仅依赖单目视频的模拟遮挡策略生成有效训练数据。此外,采用轻量级LoRA基视频扩散适配器,合成高质量、物理一致且时间连贯的视频。大量实验表明,EX-4D在物理一致性和极端视角质量上优于现有最先进方法,支持实际4D视频生成。

原文摘要 · Abstract (English)

Generating high-quality camera-controllable videos from monocular input is a challenging task, particularly under extreme viewpoint. Existing methods often struggle with geometric inconsistencies and occlusion artifacts in boundaries, leading to degraded visual quality. In this paper, we introduce EX-4D, a novel framework that addresses these challenges through a Depth Watertight Mesh representation. The representation serves as a robust geometric prior by explicitly modeling both visible and occluded regions, ensuring geometric consistency in extreme camera pose. To overcome the lack of paired multi-view datasets, we propose a simulated masking strategy that generates effective training data only from monocular videos. Additionally, a lightweight LoRA-based video diffusion adapter is employed to synthesize high-quality, physically consistent, and temporally coherent videos. Extensive experiments demonstrate that EX-4D outperforms state-of-the-art methods in terms of physical consistency and extreme-view quality, enabling practical 4D video generation.

4D视频深度网格单目生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。