arXiv:2501.17162cs.CVcs.LG2025-01ICLR被引 42

用扩散模型直接生成全景图,六面图像统一处理更高效。

CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation

  • 将全景图拆成六个面,用多视角扩散模型同步生成。
  • 无需对应注意力层,仍可生成高质量、高分辨率全景图。
  • 支持文本精细控制,适合创意设计与虚拟现实应用。

我们提出一种从文本或图像生成360°全景图的新方法。该方法利用多视角扩散模型,联合生成立方体贴图的六个面。不同于以往依赖等距投影或自回归生成的方法,本方法将每个面视为标准透视图像,简化生成流程并可直接使用现有多视角扩散模型。实验表明,这些模型无需对应感知注意力层即可生成高质量立方体贴图。该方法支持细粒度文本控制,能生成高分辨率全景图,并在训练集外具有良好的泛化能力,定性与定量指标均达到当前最优水平。

原文摘要 · Abstract (English)

We introduce a novel method for generating 360° panoramas from text prompts or images. Our approach leverages recent advances in 3D generation by employing multi-view diffusion models to jointly synthesize the six faces of a cubemap. Unlike previous methods that rely on processing equirectangular projections or autoregressive generation, our method treats each face as a standard perspective image, simplifying the generation process and enabling the use of existing multi-view diffusion models. We demonstrate that these models can be adapted to produce high-quality cubemaps without requiring correspondence-aware attention layers. Our model allows for fine-grained text control, generates high resolution panorama images and generalizes well beyond its training set, whilst achieving state-of-the-art results, both qualitatively and quantitatively. Project page: https://cubediff.github.io/

全景生成扩散模型多视角生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。