用预训练模型两阶段生成全景图,训练时间缩短至4天。
2S-ODIS: Two-Stage Omni-Directional Image Synthesis by Geometric Distortion Correction

- 先用预训练VQGAN生成粗略全景图,再融合局部图像精细修正畸变
- 相比之前方法训练时间从14天减至4天,图像质量更高
- 适合需要快速生成高质量全景图的VR与社交平台应用
全景图像在虚拟现实和社交媒体中应用日益广泛,但其获取需专用相机,数量远少于普通视角图像。现有基于生成对抗网络的方法虽可合成全景图,但存在训练不稳定、耗时长等问题。为此,本文提出2S-ODIS(两阶段全景图像合成)方法,利用在ImageNet等大规模普通视角图像数据集上预训练的VQGAN模型,无需微调即可生成高质量全景图像,并显著缩短训练时间。由于预训练模型未考虑等距柱状投影(ERP)下的畸变,该方法采用两阶段结构:第一阶段在ERP下生成全局粗略图像;第二阶段通过融合高分辨率局部普通视角图像,补偿ERP畸变。实验表明,2S-ODIS将训练时间从OmniDreamer的14天降至4天,同时提升图像质量。
原文摘要 · Abstract (English)
Omni-directional images have been increasingly used in various applications, including virtual reality and SNS (Social Networking Services). However, their availability is comparatively limited in contrast to normal field of view (NFoV) images, since specialized cameras are required to take omni-directional images. Consequently, several methods have been proposed based on generative adversarial networks (GAN) to synthesize omni-directional images, but these approaches have shown difficulties in training of the models, due to instability and/or significant time consumption in the training. To address these problems, this paper proposes a novel omni-directional image synthesis method, 2S-ODIS (Two-Stage Omni-Directional Image Synthesis), which generated high-quality omni-directional images but drastically reduced the training time. This was realized by utilizing the VQGAN (Vector Quantized GAN) model pre-trained on a large-scale NFoV image database such as ImageNet without fine-tuning. Since this pre-trained model does not represent distortions of omni-directional images in the equi-rectangular projection (ERP), it cannot be applied directly to the omni-directional image synthesis in ERP. Therefore, two-stage structure was adopted to first create a global coarse image in ERP and then refine the image by integrating multiple local NFoV images in the higher resolution to compensate the distortions in ERP, both of which are based on the pre-trained VQGAN model. As a result, the proposed method, 2S-ODIS, achieved the reduction of the training time from 14 days in OmniDreamer to four days in higher image quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。