arXiv:2505.16862cs.CV2025-05NeurIPS被引 8

用自回归模型统一生成全景图,解决文本和图像条件下的连贯性问题。

Conditional Panoramic Image Generation via Masked Autoregressive Modeling

  • 采用掩码自回归建模,避开扩散模型对独立同分布的依赖。
  • 在文本生成与全景外扩任务中均表现优异,支持跨任务无缝切换。
  • 引入循环填充与一致性对齐,提升全景图空间连续性与质量。

近期全景图像生成研究揭示了现有方法的两大局限:其一,多数基于扩散模型,而扩散模型在等距圆柱投影(ERP)全景图上因球面映射破坏独立同分布高斯噪声假设,表现不佳;其二,文本条件生成(文本到全景图)与图像条件生成(全景外扩)常被当作独立任务处理,依赖不同架构和专用数据。本文提出统一框架——全景自回归模型(PAR),采用掩码自回归建模,规避 i.i.d. 假设约束,并将文本与图像条件整合至同一架构中,实现任务间的无缝生成。为应对生成模型固有的不连续性,引入循环填充以增强空间连贯性,并提出一致性对齐策略提升生成质量。大量实验表明,该方法在文本生成与全景外扩任务中均达到竞争力水平,展现出良好的可扩展性与泛化能力。

原文摘要 · Abstract (English)

Recent progress in panoramic image generation has underscored two critical limitations in existing approaches. First, most methods are built upon diffusion models, which are inherently ill-suited for equirectangular projection (ERP) panoramas due to the violation of the identically and independently distributed (i.i.d.) Gaussian noise assumption caused by their spherical mapping. Second, these methods often treat text-conditioned generation (text-to-panorama) and image-conditioned generation (panorama outpainting) as separate tasks, relying on distinct architectures and task-specific data. In this work, we propose a unified framework, Panoramic AutoRegressive model (PAR), which leverages masked autoregressive modeling to address these challenges. PAR avoids the i.i.d. assumption constraint and integrates text and image conditioning into a cohesive architecture, enabling seamless generation across tasks. To address the inherent discontinuity in existing generative models, we introduce circular padding to enhance spatial coherence and propose a consistency alignment strategy to improve generation quality. Extensive experiments demonstrate competitive performance in text-to-image generation and panorama outpainting tasks while showcasing promising scalability and generalization capabilities.

全景生成自回归模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。