arXiv:2511.13121cs.CV2025-11被引 2

用点条件扩散模型实现稀疏视角下的近景新视角合成

CloseUpShot: Close-up Novel View Synthesis from Sparse-views via Point-conditioned Diffusion Model

  • 通过点条件扩散模型,解决近景下视角稀疏导致的细节丢失问题
  • 在多个数据集上显著优于现有方法,近景合成质量提升明显
  • 适合需要高精度近景重建的应用,如文物数字化、医疗成像

从稀疏视角重建3D场景并合成新视角是一项极具挑战性的任务。近年来视频扩散模型展现出强大的时序推理能力,为稀疏视角下的重建质量提升提供了新思路。然而,现有方法主要针对视角变化较小的情况,在近景场景中因输入信息严重不足,难以捕捉精细细节。本文提出一种基于扩散的框架CloseUpShot,通过点条件视频扩散模型实现稀疏视角下的近景新视角合成。我们发现像素对齐条件在近景设置下存在严重稀疏性和背景泄露问题。为此,提出分层对齐与遮挡感知噪声抑制机制,提升了扩散模型条件图像的质量与完整性。此外,引入全局结构引导,利用密集融合点云为扩散过程提供一致的几何上下文,弥补稀疏条件输入中缺乏全局一致性3D约束的问题。大量实验表明,该方法在多个数据集上均优于现有方法,尤其在近景新视角合成任务中表现突出,充分验证了设计的有效性。

原文摘要 · Abstract (English)

Reconstructing 3D scenes and synthesizing novel views from sparse input views is a highly challenging task. Recent advances in video diffusion models have demonstrated strong temporal reasoning capabilities, making them a promising tool for enhancing reconstruction quality under sparse-view settings. However, existing approaches are primarily designed for modest viewpoint variations, which struggle in capturing fine-grained details in close-up scenarios since input information is severely limited. In this paper, we present a diffusion-based framework, called CloseUpShot, for close-up novel view synthesis from sparse inputs via point-conditioned video diffusion. Specifically, we observe that pixel-warping conditioning suffers from severe sparsity and background leakage in close-up settings. To address this, we propose hierarchical warping and occlusion-aware noise suppression, enhancing the quality and completeness of the conditioning images for the video diffusion model. Furthermore, we introduce global structure guidance, which leverages a dense fused point cloud to provide consistent geometric context to the diffusion process, to compensate for the lack of globally consistent 3D constraints in sparse conditioning inputs. Extensive experiments on multiple datasets demonstrate that our method outperforms existing approaches, especially in close-up novel view synthesis, clearly validating the effectiveness of our design.

新视角合成扩散模型近景重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。