视频与4D内容相互优化,生成更逼真的动态三维场景。
Video4DGen: Enhancing Video and 4D Generation through Mutual Optimization
- 用动态高斯表面元(DGS)建模时变形变,实现时空一致的4D生成。
- 多视频对齐与姿态引导采样,提升跨视角、跨时间的结构一致性。
- 支持新视角视频生成,适合虚拟现实与动画制作人员使用。
4D(即序列3D)生成技术的发展为虚拟体验开辟了新可能,用户可从任意视角探索动态物体或角色。与此同时,视频生成模型因其能产出逼真且富有想象力的帧而备受关注,且表现出强3D一致性,暗示其具备世界模拟潜力。本文提出Video4DGen框架,能够从单个或多个生成视频中生成4D表示,并实现4D引导下的视频生成。该框架采用动态高斯表面元(DGS)表示4D输出,通过优化时变变形函数,将静态高斯表面元转换为动态形态。设计了变形状态几何正则化与精炼机制,以保持结构完整性和细节特征。为从多视频中进行4D生成并捕捉空间、时间与姿态维度的表征,引入多视频对齐、根姿态优化和姿态引导帧采样策略。连续变形场的使用实现了对每帧中姿态、运动与形变的精确刻画。此外,为提升整体保真度,Video4DGen利用4D内容指导新视角视频生成,结合置信度过滤的DGS增强生成序列质量。该框架在虚拟现实、动画等领域具有广泛应用潜力。
原文摘要 · Abstract (English)
The advancement of 4D (i.e., sequential 3D) generation opens up new possibilities for lifelike experiences in various applications, where users can explore dynamic objects or characters from any viewpoint. Meanwhile, video generative models are receiving particular attention given their ability to produce realistic and imaginative frames. These models are also observed to exhibit strong 3D consistency, indicating the potential to act as world simulators. In this work, we present Video4DGen, a novel framework that excels in generating 4D representations from single or multiple generated videos as well as generating 4D-guided videos. This framework is pivotal for creating high-fidelity virtual contents that maintain both spatial and temporal coherence. The 4D outputs generated by Video4DGen are represented using our proposed Dynamic Gaussian Surfels (DGS), which optimizes time-varying warping functions to transform Gaussian surfels (surface elements) from a static state to a dynamically warped state. We design warped-state geometric regularization and refinements on Gaussian surfels, to preserve the structural integrity and fine-grained appearance details. To perform 4D generation from multiple videos and capture representation across spatial, temporal, and pose dimensions, we design multi-video alignment, root pose optimization, and pose-guided frame sampling strategies. The leveraging of continuous warping fields also enables a precise depiction of pose, motion, and deformation over per-video frames. Further, to improve the overall fidelity from the observation of all camera poses, Video4DGen performs novel-view video generation guided by the 4D content, with the proposed confidence-filtered DGS to enhance the quality of generated sequences. With the ability of 4D and video generation, Video4DGen offers a powerful tool for applications in virtual reality, animation, and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。