从单图生成可360°旋转缩放的高清3D场景
FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis
- 用视频到视频扩散模型生成新视角图像
- 通过渐进式扩展构建完整3D场景,支持大角度相机变化
- 适合需要高质量自由视角合成的应用场景
从单张图像生成支持灵活视角(如360°旋转和缩放)的3D场景极具挑战,主要因缺乏三维数据。为此,我们提出FlexWorld框架,包含两个核心组件:(1) 基于强大视频到视频(V2V)扩散模型,从粗略场景渲染的不完整输入中生成高质量新视角图像;(2) 渐进式扩展过程,逐步生成新3D内容并以几何感知方式融合至全局场景。该V2V模型利用预训练视频模型与精确深度估计训练对,在大相机姿态变化下仍能生成高质量图像。在此基础上,FlexWorld通过几何感知融合逐步构建完整3D场景。大量实验表明,该方法在多个主流数据集上优于现有最先进方法,在多种评价指标下均取得更优视觉质量。定性结果显示,FlexWorld可生成支持360°旋转和缩放的高保真场景。
原文摘要 · Abstract (English)
Generating flexible-view 3D scenes, including 360° rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework consisting of two key components: (1) a strong video-to-video (V2V) diffusion model to generate high-quality novel view images from incomplete input rendered from a coarse scene, and (2) a progressive expansion process to construct a complete 3D scene. In particular, leveraging an advanced pre-trained video model and accurate depth-estimated training pairs, our V2V model can generate novel views under large camera pose variations. Building upon it, FlexWorld progressively generates new 3D content and integrates it into the global scene through geometry-aware scene fusion. Extensive experiments demonstrate the effectiveness of FlexWorld in generating high-quality novel view videos and flexible-view 3D scenes from single images, achieving superior visual quality under multiple popular metrics and datasets compared to existing state-of-the-art methods. Qualitatively, we highlight that FlexWorld can generate high-fidelity scenes with flexible views like 360° rotations and zooming. Project page: https://ml-gsai.github.io/FlexWorld.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。