arXiv:2512.06885cs.CVcs.AI2025-12

用统一模型同时生成文本和视角驱动的全景图,效果更优。

JoPano: Unified Panorama Generation via Joint Modeling

  • 基于DiT架构,用立方体表示融合多视角生成
  • 在FID、CLIP-FID等四项指标上达到顶尖水平
  • 适合需要高效生成高质量全景图的研究者

全景图生成近年来受到广泛关注,主要包含文本到全景图和视角到全景图两类任务。现有方法仍面临两大挑战:基于U-Net的架构限制了生成质量,且两类任务常独立处理,导致建模冗余与效率低下。为此,本文提出联合全景生成方法JoPano,将两类任务统一于DiT-based模型中。为将自然图像预训练的DiT骨干能力迁移到全景域,提出基于立方体表示的联合面适配器(Joint-Face Adapter),使预训练DiT可协同生成全景各视角。进一步采用泊松混合(Poisson Blending)减少立方体面间接缝不一致问题,并引入塞姆-SSIM和塞姆-Sobel指标定量评估接缝一致性。此外,设计条件切换机制,使单一模型支持文本与视角驱动生成。大量实验表明,JoPano在文本到全景图和视角到全景图任务中均能生成高质量全景图,在FID、CLIP-FID、IS和CLIP-Score指标上均达当前最优表现。

原文摘要 · Abstract (English)

Panorama generation has recently attracted growing interest in the research community, with two core tasks, text-to-panorama and view-to-panorama generation. However, existing methods still face two major challenges: their U-Net-based architectures constrain the visual quality of the generated panoramas, and they usually treat the two core tasks independently, which leads to modeling redundancy and inefficiency. To overcome these challenges, we propose a joint-face panorama (JoPano) generation approach that unifies the two core tasks within a DiT-based model. To transfer the rich generative capabilities of existing DiT backbones learned from natural images to the panorama domain, we propose a Joint-Face Adapter built on the cubemap representation of panoramas, which enables a pretrained DiT to jointly model and generate different views of a panorama. We further apply Poisson Blending to reduce seam inconsistencies that often appear at the boundaries between cube faces. Correspondingly, we introduce Seam-SSIM and Seam-Sobel metrics to quantitatively evaluate the seam consistency. Moreover, we propose a condition switching mechanism that unifies text-to-panorama and view-to-panorama tasks within a single model. Comprehensive experiments show that JoPano can generate high-quality panoramas for both text-to-panorama and view-to-panorama generation tasks, achieving state-of-the-art performance on FID, CLIP-FID, IS, and CLIP-Score metrics.

全景生成DiT图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。