用自回归扩散模型提升3D重建在未观测区域的生成质量与效率
ArtiFixer: Enhancing and Extending 3D Reconstruction with Auto-Regressive Diffusion Models

- 提出两阶段流程:先训练双向生成模型,再蒸馏为可单次生成数百帧的自回归模型
- 在未观测区域生成内容更一致,比现有方法提升1-3 dB PSNR
- 适合需要高保真3D重建且关注未观测区域泛化的研究者
基于场景的优化方法如3D高斯点云(3D Gaussian Splatting)在新视角合成上表现优异,但在观测不足区域外推能力差。利用生成先验修复这些区域的方法虽有潜力,但存在两大缺陷:一是可扩展性差,现有方法使用图像扩散模型或双向视频模型,单次生成视图数量有限,需昂贵的迭代蒸馏过程以保证一致性;二是生成质量不佳,先前生成器输出常与已有场景内容不一致,完全未观测区域甚至无法生成。为此,我们提出两阶段管道,利用两个关键洞察:首先,训练一个强大的双向生成模型,结合新颖的透明度混合策略,既保持与已有观测的一致性,又保留生成新内容的能力;其次,将其蒸馏为因果自回归模型,可在单次推理中生成数百帧。该模型可直接生成新视角,或作为伪监督信号高效优化底层3D表示。我们在多个基准数据集上全面评估,证明其能在现有方法完全失效的情况下生成合理重建,在常见数据集上优于所有基线,领先于最先进方法1-3 dB PSNR。
原文摘要 · Abstract (English)
Per-scene optimization methods such as 3D Gaussian Splatting provide state-of-the-art novel view synthesis quality but extrapolate poorly to under-observed areas. Methods that leverage generative priors to correct artifacts in these areas hold promise but currently suffer from two shortcomings. The first is scalability, as existing methods use image diffusion models or bidirectional video models that are limited in the number of views they can generate in a single pass (and thus require a costly iterative distillation process for consistency). The second is quality itself, as generators used in prior work tend to produce outputs that are inconsistent with existing scene content and fail entirely in completely unobserved regions. To solve these, we propose a two-stage pipeline that leverages two key insights. First, we train a powerful bidirectional generative model with a novel opacity mixing strategy that encourages consistency with existing observations while retaining the model's ability to extrapolate novel content in unseen areas. Second, we distill it into a causal auto-regressive model that generates hundreds of frames in a single pass. This model can directly produce novel views or serve as pseudo-supervision to improve the underlying 3D representation in a simple and highly efficient manner. We evaluate our method extensively and demonstrate that it can generate plausible reconstructions in scenarios where existing approaches fail completely. When measured on commonly benchmarked datasets, we outperform all existing baselines by a wide margin, exceeding prior state-of-the-art methods by 1-3 dB PSNR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。