arXiv:2508.06392cs.CV2025-08ICCV被引 1

用对抗蒸馏让视频扩散模型4步生成新视角,速度提升90%以上

FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation

  • 用GAN和软KL散度将多步视频扩散模型压缩为4步学生模型
  • 在真实数据集上保持画质的同时,采样时间减少90%以上
  • 特别适合稀疏输入视角的3D重建,大幅降低计算开销

近期3D重建进展实现了从密集图像捕捉生成逼真3D模型,但稀疏视角仍会导致未见区域出现伪影。现有方法利用视频扩散模型(VDMs)生成密集观测以填补稀疏视角下的空缺。然而,这些方法使用VDMs时采样速度慢。本文提出FVGen框架,通过对抗性视频扩散蒸馏,仅需4步采样即可实现快速新视角合成。我们提出一种新型蒸馏方法,将多步去噪教师模型通过生成对抗网络(GAN)和软化反向KL散度最小化,压缩为少步去噪学生模型。在真实世界数据集上的大量实验表明,相比之前方法,本框架在生成相同数量新视角时,视觉质量相当甚至更优,同时采样时间减少超过90%。FVGen显著提升了下游重建任务的时间效率,尤其适用于输入视角稀疏(超过2个)的情况,此时预训练的VDM需要多次运行以获得更好空间覆盖。

原文摘要 · Abstract (English)

Recent progress in 3D reconstruction has enabled realistic 3D models from dense image captures, yet challenges persist with sparse views, often leading to artifacts in unseen areas. Recent works leverage Video Diffusion Models (VDMs) to generate dense observations, filling the gaps when only sparse views are available for 3D reconstruction tasks. A significant limitation of these methods is their slow sampling speed when using VDMs. In this paper, we present FVGen, a novel framework that addresses this challenge by enabling fast novel view synthesis using VDMs in as few as four sampling steps. We propose a novel video diffusion model distillation method that distills a multi-step denoising teacher model into a few-step denoising student model using Generative Adversarial Networks (GANs) and softened reverse KL-divergence minimization. Extensive experiments on real-world datasets show that, compared to previous works, our framework generates the same number of novel views with similar (or even better) visual quality while reducing sampling time by more than 90%. FVGen significantly improves time efficiency for downstream reconstruction tasks, particularly when working with sparse input views (more than 2) where pre-trained VDMs need to be run multiple times to achieve better spatial coverage.

视频生成扩散模型3D重建加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。