arXiv:2502.11697cs.CV2025-02ICLR被引 18

用多视角像素流生成高质量动态4D内容,解决时空一致性难题。

MVTokenFlow: High-quality 4D Content Generation using Multiview Token Flow

  • 通过多视角扩散模型生成分时多视角图像,构建粗略4D场。
  • 利用2D光流引导重生成图像,提升时间一致性并复用视觉特征。
  • 适合需要高保真动态3D视频生成的研究者与开发者。

本文提出MVTokenFlow,用于从单目视频生成高质量4D内容。当前基于视频扩散模型和多视角扩散模型的生成方法虽可创建视频或3D模型,但将其扩展至动态4D内容生成仍面临时空一致性挑战。MVTokenFlow首先使用多视角扩散模型在不同时间步生成多视角图像,实现跨视角的空间一致性,从而重建合理的粗略4D场。随后,该方法以渲染的2D光流为引导,对所有多视角图像进行重生成,2D光流有效关联不同时刻的像素,并通过重用生成过程中的视觉令牌来增强时间一致性。最终生成的图像具备良好的时空一致性,并用于精细化优化粗略4D场,获得高质量4D场。实验表明,本方法设计有效,生成质量显著优于基线方法。

原文摘要 · Abstract (English)

In this paper, we present MVTokenFlow for high-quality 4D content creation from monocular videos. Recent advancements in generative models such as video diffusion models and multiview diffusion models enable us to create videos or 3D models. However, extending these generative models for dynamic 4D content creation is still a challenging task that requires the generated content to be consistent spatially and temporally. To address this challenge, MVTokenFlow utilizes the multiview diffusion model to generate multiview images on different timesteps, which attains spatial consistency across different viewpoints and allows us to reconstruct a reasonable coarse 4D field. Then, MVTokenFlow further regenerates all the multiview images using the rendered 2D flows as guidance. The 2D flows effectively associate pixels from different timesteps and improve the temporal consistency by reusing tokens in the regeneration process. Finally, the regenerated images are spatiotemporally consistent and utilized to refine the coarse 4D field to get a high-quality 4D field. Experiments demonstrate the effectiveness of our design and show significantly improved quality than baseline methods.

4D生成扩散模型多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。