一个模型搞定全景图生成与编辑,解决传统方法不适用问题。
Omni$^2$: Unifying Omnidirectional Image Generation and Editing in an Omni Model
- 用统一模型处理多种全景图生成与编辑任务
- 构建包含6万+数据的首个全景图多任务数据集
- 适合虚拟现实、增强现实领域研究人员使用
360°全景图像(ODI)近年来受到广泛关注,广泛应用于虚拟现实(VR)和增强现实(AR)场景。但捕获此类图像成本高且需专用设备,因此合成技术愈发重要。现有2D图像生成与编辑方法在处理全景图像时因格式特殊、视角广而表现不佳。为此,我们构建了首个综合性全景图像生成与编辑数据集Any2Omni,包含超过6万条训练数据,覆盖多种输入条件及最多9类任务。基于该数据集,提出统一模型Omni²,可在一个模型中完成多种全景图像生成与编辑任务。大量实验验证了Omni²在生成与编辑任务上的优越性。相关数据集与模型已开源:https://github.com/IntMeGroup/Omni2。
原文摘要 · Abstract (English)
$360^{\circ}$ omnidirectional images (ODIs) have gained considerable attention recently, and are widely used in various virtual reality (VR) and augmented reality (AR) applications. However, capturing such images is expensive and requires specialized equipment, making ODI synthesis increasingly important. While common 2D image generation and editing methods are rapidly advancing, these models struggle to deliver satisfactory results when generating or editing ODIs due to the unique format and broad 360$^{\circ}$ Field-of-View (FoV) of ODIs. To bridge this gap, we construct \textbf{\textit{Any2Omni}}, the first comprehensive ODI generation-editing dataset comprises 60,000+ training data covering diverse input conditions and up to 9 ODI generation and editing tasks. Built upon Any2Omni, we propose an \textbf{\underline{Omni}} model for \textbf{\underline{Omni}}-directional image generation and editing (\textbf{\textit{Omni$^2$}}), with the capability of handling various ODI generation and editing tasks under diverse input conditions using one model. Extensive experiments demonstrate the superiority and effectiveness of the proposed Omni$^2$ model for both the ODI generation and editing tasks. Both the Any2Omni dataset and the Omni$^2$ model are publicly available at: https://github.com/IntMeGroup/Omni2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。