让全景图编辑更真实,自动构建高质量数据集
SE360: Semantic Edit in 360$^\circ$ Panoramas via Hierarchical Data Construction
- 分层自动构建数据:用视觉语言模型生成无标注全景图的语义-几何一致数据对
- 编辑效果更好:在等距圆柱投影和透视视图中均实现更高视觉质量与语义准确率
- 适合做全景内容创作的人:如虚拟现实、元宇宙场景编辑者
尽管基于指令的图像编辑正在兴起,但将其扩展到360°全景图仍面临额外挑战。现有方法在等距圆柱投影(ERP)和透视视图中常产生不合理结果。为此,我们提出SE360,一种用于360°全景图多条件引导物体编辑的新框架。核心是无需人工干预的粗到细自主数据生成流程,利用视觉语言模型(VLM)和自适应投影调整进行分层分析,确保物体及其物理上下文的整体分割。生成的数据对在语义上合理且几何上一致,即使来自未标注全景图。此外,我们引入低成本的两阶段数据精炼策略,提升数据真实感并减轻模型过拟合导致的伪影。基于构建的数据集,训练了一个基于Transformer的扩散模型,支持文本、掩码或参考图像引导的灵活物体编辑。实验表明,该方法在视觉质量和语义准确性上优于现有方法。
原文摘要 · Abstract (English)
While instruction-based image editing is emerging, extending it to 360$^\circ$ panoramas introduces additional challenges. Existing methods often produce implausible results in both equirectangular projections (ERP) and perspective views. To address these limitations, we propose SE360, a novel framework for multi-condition guided object editing in 360$^\circ$ panoramas. At its core is a novel coarse-to-fine autonomous data generation pipeline without manual intervention. This pipeline leverages a Vision-Language Model (VLM) and adaptive projection adjustment for hierarchical analysis, ensuring the holistic segmentation of objects and their physical context. The resulting data pairs are both semantically meaningful and geometrically consistent, even when sourced from unlabeled panoramas. Furthermore, we introduce a cost-effective, two-stage data refinement strategy to improve data realism and mitigate model overfitting to erase artifacts. Based on the constructed dataset, we train a Transformer-based diffusion model to allow flexible object editing guided by text, mask, or reference image in 360$^\circ$ panoramas. Our experiments demonstrate that our method outperforms existing methods in both visual quality and semantic accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。