arXiv:2511.08823cs.CVcs.AI2025-11被引 1

用单图生成真实场景新视角,支持多样化输出。

DT-NVS: Diffusion Transformers for Novel View Synthesis

  • 基于Transformer的3D扩散模型,直接从图像学习3D表示。
  • 在真实未对齐视频数据上训练,实现跨视角泛化生成。
  • 创新相机条件设计与训练范式,适合日常场景应用。

从单张图像生成自然场景(包括室内外)的新视角是一个未充分探索的问题,尽管这可视为以物体为中心的视角合成的自然延伸。现有基于扩散的方法主要关注真实场景中微小相机运动或仅处理非自然的物体中心场景,限制了其在真实世界中的应用。本文提出一种3D感知扩散模型DT-NVS,该模型在大规模真实世界、多类别、未对齐且随意拍摄的日常场景视频数据集上,仅使用图像级损失进行训练。通过改进Transformer与自注意力架构,实现图像到3D表示的转换,并引入新颖的相机条件策略,使模型可在真实世界未对齐数据上训练。此外,我们提出一种新的训练范式,交换参考帧在条件图像与噪声输入之间的角色。我们在单图像泛化视角合成任务上评估方法,结果优于现有的3D感知扩散模型和确定性方法,同时生成多样化的输出。

原文摘要 · Abstract (English)

Generating novel views of a natural scene, e.g., every-day scenes both indoors and outdoors, from a single view is an under-explored problem, even though it is an organic extension to the object-centric novel view synthesis. Existing diffusion-based approaches focus rather on small camera movements in real scenes or only consider unnatural object-centric scenes, limiting their potential applications in real-world settings. In this paper we move away from these constrained regimes and propose a 3D diffusion model trained with image-only losses on a large-scale dataset of real-world, multi-category, unaligned, and casually acquired videos of everyday scenes. We propose DT-NVS, a 3D-aware diffusion model for generalized novel view synthesis that exploits a transformer-based architecture backbone. We make significant contributions to transformer and self-attention architectures to translate images to 3d representations, and novel camera conditioning strategies to allow training on real-world unaligned datasets. In addition, we introduce a novel training paradigm swapping the role of reference frame between the conditioning image and the sampled noisy input. We evaluate our approach on the 3D task of generalized novel view synthesis from a single input image and show improvements over state-of-the-art 3D aware diffusion models and deterministic approaches, while generating diverse outputs.

3D生成扩散模型视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。