通过几何感知预训练,提升上下文全景生成的准确性与一致性。
Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

- 分两阶段框架:先几何感知预训练,再任务微调。
- 100万高质量全景样本数据集,支持多种编辑任务。
- 融合深度生成与环形填充,显著提升全景几何一致性。
本文提出Canvas360,一种两阶段的上下文全景生成框架,结合几何感知预训练与下游任务微调。针对缺乏大规模高质量上下文全景训练数据的问题,构建了包含100万组高质量配对样本的Canvas360Dataset,涵盖风格迁移、修复、外扩和编辑等任务,实现多样化上下文生成场景的有效监督。在建模层面,通过并行深度生成、速度导向环形填充及相似性损失正则化,使模型学习几何感知表示,捕捉物体畸变细节,提升几何一致性和全局连贯性。得益于强大的全景先验,Canvas360实现统一的上下文全景生成框架,通过标记级拼接支持多种下游任务,在任务覆盖和建模灵活性上超越现有方法。大量实验表明,该模型显著提升全景图像保真度,在专属的FAED指标上表现优异,且在各项定量评估中达到竞争性或领先水平。
原文摘要 · Abstract (English)
In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning. To address the lack of large-scale, high-quality training data tailored to in-context panoramic tasks, we propose Canvas360Dataset, a collection of 1M high-quality paired panoramic samples for style transfer, inpainting, outpainting, and editing, enabling effective supervision across diverse in-context generation scenarios. On the modeling side, Canvas360 enhances text-to-panorama generation through parallel depth generation, velocity circular padding, and similarity loss regularization, enabling the model to learn geometry-aware representations, capture object distortion details, and improve geometric consistency and global coherence. Furthermore, empowered by strong panoramic priors, Canvas360 enables a unified in-context panoramic generation framework that supports diverse downstream tasks via token-level concatenation, surpassing prior methods in both task coverage and modeling flexibility. Extensive experiments show that Canvas360 improves panoramic image fidelity, achieving particularly strong performance on the panorama-specific FAED metric and competitive or leading results across the reported quantitative evaluations. More information can be found on our project page: https://zry000.github.io/Canvas360/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。