用视觉变压器+正弦网络提升气候模型图像超分辨率重建质量
ViSIR: Vision Transformer Single Image Reconstruction Method for Earth System Models
- 将ViT与SIREN结合,解决超分任务中的频谱偏差问题
- 在三个数据集上平均比SIREN高8.34dB PSNR
- 适合需要高精度气候数据重建的研究者使用
地球系统模型(ESMs)整合大气、海洋、陆地、冰和生物圈的相互作用,以估计各种条件下的区域和全球气候状态。由于模型高度复杂,常采用深度神经网络架构来建模复杂性并存储降采样数据。本文提出视觉变压器正弦表示网络(ViSIR),用于改进地球系统模型数据的单图像超分辨率(SR)重建任务。ViSIR将视觉变压器(ViT)的超分能力与正弦表示网络(SIREN)的高频细节保留特性相结合,以应对超分任务中常见的频谱偏差问题。实验结果显示,与SRCNN相比,ViSIR平均提升2.16 dB;与ViT相比提升6.29 dB;与SIREN相比提升8.34 dB;与SR-GAN相比提升7.93 dB PSNR。在均方误差(MSE)、峰值信噪比(PSNR)和结构相似性指数(SSIM)三项指标上,所提方法均优于现有先进方法。
原文摘要 · Abstract (English)
Purpose: Earth system models (ESMs) integrate the interactions of the atmosphere, ocean, land, ice, and biosphere to estimate the state of regional and global climate under a wide variety of conditions. The ESMs are highly complex; thus, deep neural network architectures are used to model the complexity and store the down-sampled data. This paper proposes the Vision Transformer Sinusoidal Representation Networks (ViSIR) to improve the ESM data's single image SR (SR) reconstruction task. Methods: ViSIR combines the SR capability of Vision Transformers (ViT) with the high-frequency detail preservation of the Sinusoidal Representation Network (SIREN) to address the spectral bias observed in SR tasks. Results: The ViSIR outperforms SRCNN by 2.16 db, ViT by 6.29 dB, SIREN by 8.34 dB, and SR-Generative Adversarial (SRGANs) by 7.93 dB PSNR on average for three different measurements. Conclusion: The proposed ViSIR is evaluated and compared with state-of-the-art methods. The results show that the proposed algorithm is outperforming other methods in terms of Mean Square Error(MSE), Peak-Signal-to-Noise-Ratio(PSNR), and Structural Similarity Index Measure(SSIM).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。