arXiv:2504.01956cs.CV2025-04CVPR被引 32

用视频扩散模型一键生成3D场景,速度快且结构更真实

VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step

  • 通过3D感知的跳跃流蒸馏,跳过冗余时间信息
  • 动态去噪网络自适应选择最优推理步数,提升效率
  • 适合需要快速高质量3D重建的视觉应用开发者

从稀疏视角恢复3D场景是固有病态问题。传统方法依赖几何正则化或前馈确定性模型,但在输入视角重叠度低、视觉信息不足时性能下降明显。近期视频生成模型展现出生成具合理3D结构视频的潜力。基于预训练视频扩散模型,已有研究探索利用视频生成先验从稀疏视图构建3D场景,但存在推理慢、缺乏3D约束的问题,导致效率低下和重建伪影。本文提出VideoScene,通过蒸馏视频扩散模型实现一步生成3D场景,旨在构建高效可靠的视频到3D工具。设计3D感知跳跃流蒸馏策略,跳过耗时冗余信息;训练动态去噪策略网络,自适应确定推理中的最优跳跃时间步。大量实验表明,VideoScene在速度与重建质量上均优于现有视频扩散模型,展现出未来视频转3D应用的巨大潜力。

原文摘要 · Abstract (English)

Recovering 3D scenes from sparse views is a challenging task due to its inherent ill-posed problem. Conventional methods have developed specialized solutions (e.g., geometry regularization or feed-forward deterministic model) to mitigate the issue. However, they still suffer from performance degradation by minimal overlap across input views with insufficient visual information. Fortunately, recent video generative models show promise in addressing this challenge as they are capable of generating video clips with plausible 3D structures. Powered by large pretrained video diffusion models, some pioneering research start to explore the potential of video generative prior and create 3D scenes from sparse views. Despite impressive improvements, they are limited by slow inference time and the lack of 3D constraint, leading to inefficiencies and reconstruction artifacts that do not align with real-world geometry structure. In this paper, we propose VideoScene to distill the video diffusion model to generate 3D scenes in one step, aiming to build an efficient and effective tool to bridge the gap from video to 3D. Specifically, we design a 3D-aware leap flow distillation strategy to leap over time-consuming redundant information and train a dynamic denoising policy network to adaptively determine the optimal leap timestep during inference. Extensive experiments demonstrate that our VideoScene achieves faster and superior 3D scene generation results than previous video diffusion models, highlighting its potential as an efficient tool for future video to 3D applications. Project Page: https://hanyang-21.github.io/VideoScene

视频生成3D重建扩散模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。