用全景适配的Transformer模型,提升全景HDR光照估计质量。
360U-Former: HDR Illumination Estimation with Panoramic Adapted Vision Transformers
- 基于U-Net结构的视觉Transformer,适配等距圆柱投影格式。
- 生成无接缝、无极点畸变的高质量全景HDR图像。
- 首个纯Transformer架构用于光照估计,适合虚拟现实与3D渲染场景。
现有光照估计方法多关注提升分辨率和生成纹理的质量与多样性,但较少针对图像光照中常用的等距圆柱投影(ERP)格式优化神经网络架构。因此,生成的高动态范围图像(HDRI)常在侧边出现接缝,极区纹理或物体发生扭曲。为此,我们提出一种新型架构360U-Former,基于具有U-Net结构的视觉Transformer,借鉴PanoSWIN的适配ERP的移位窗口注意力机制。据我们所知,这是首个在光照估计领域完全采用视觉Transformer的模型。我们以生成对抗网络(GAN)方式训练360U-Former,从有限视场低动态范围图像(LDRI)生成HDRI。在当前光照估计评估协议与数据集上验证表明,本方法优于现有及最先进方法,且避免了使用ERP格式常见的伪影。
原文摘要 · Abstract (English)
Recent illumination estimation methods have focused on enhancing the resolution and improving the quality and diversity of the generated textures. However, few have explored tailoring the neural network architecture to the Equirectangular Panorama (ERP) format utilised in image-based lighting. Consequently, high dynamic range images (HDRI) results usually exhibit a seam at the side borders and textures or objects that are warped at the poles. To address this shortcoming we propose a novel architecture, 360U-Former, based on a U-Net style Vision-Transformer which leverages the work of PanoSWIN, an adapted shifted window attention tailored to the ERP format. To the best of our knowledge, this is the first purely Vision-Transformer model used in the field of illumination estimation. We train 360U-Former as a GAN to generate HDRI from a limited field of view low dynamic range image (LDRI). We evaluate our method using current illumination estimation evaluation protocols and datasets, demonstrating that our approach outperforms existing and state-of-the-art methods without the artefacts typically associated with the use of the ERP format.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。