用对齐表征的生成模型提升稀疏照片下的3D场景新视角合成质量。
BetterScene: 3D Scene Synthesis with Representation-Aligned Generative Model
- 在SVD框架中改进VAE模块,引入时序等变与视觉基础对齐正则化。
- 在DL3DV-10K数据集上实现更一致、无伪影的新视角渲染效果。
- 适合做高质量3D场景重建的科研与工业应用,尤其关注细节一致性。
我们提出BetterScene,一种利用极稀疏、无约束照片提升多样化真实场景新视角合成质量的方法。该方法采用预训练于数十亿帧的稳定视频扩散(SVD)模型作为骨干,旨在缓解合成中的伪影并恢复视图一致的细节。现有基于扩散的方法通常仅微调UNet模块,冻结其他组件,即便加入深度或语义等几何感知正则化,仍存在细节不一致和伪影问题。为此,我们研究了扩散模型的潜在空间,提出两个新组件:(1)时序等变正则化,(2)视觉基础模型对齐表示,均应用于SVD流程中的变分自编码器(VAE)。BetterScene结合前馈式3D高斯溅射(3DGS)模型,将特征渲染作为SVD增强器输入,生成连续、无伪影且一致的新视角。我们在具有挑战性的DL3DV-10K数据集上评估,表现优于现有最先进方法。
原文摘要 · Abstract (English)
We present BetterScene, an approach to enhance novel view synthesis (NVS) quality for diverse real-world scenes using extremely sparse, unconstrained photos. BetterScene leverages the production-ready Stable Video Diffusion (SVD) model pretrained on billions of frames as a strong backbone, aiming to mitigate artifacts and recover view-consistent details at inference time. Conventional methods have developed similar diffusion-based solutions to address these challenges of novel view synthesis. Despite significant improvements, these methods typically rely on off-the-shelf pretrained diffusion priors and fine-tune only the UNet module while keeping other components frozen, which still leads to inconsistent details and artifacts even when incorporating geometry-aware regularizations like depth or semantic conditions. To address this, we investigate the latent space of the diffusion model and introduce two components: (1) temporal equivariance regularization and (2) vision foundation model-aligned representation, both applied to the variational autoencoder (VAE) module within the SVD pipeline. BetterScene integrates a feed-forward 3D Gaussian Splatting (3DGS) model to render features as inputs for the SVD enhancer and generate continuous, artifact-free, consistent novel views. We evaluate on the challenging DL3DV-10K dataset and demonstrate superior performance compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。