提出两种频域感知视觉变换器,提升气候模型超分辨率细节还原能力
Frequency-Aware Vision Transformers for High-Fidelity Super-Resolution of Earth System Models
- 用正弦激活和傅里叶滤波增强变压器对高频信息的捕捉能力
- 在E3SM-HR数据集上,峰值信噪比最高提升2.6 dB,结构相似性更优
- 适合需要高精度气候模拟输出的研究者,尤其关注细粒度空间结构
超分辨率技术在提升地球系统模型输出的空间保真度方面具有重要作用,有助于从粗分辨率模拟中恢复对气候科学至关重要的细尺度结构。然而,传统的深度超分辨率方法(包括卷积和基于Transformer的模型)普遍存在频谱偏差问题,更易重建低频内容而忽略有价值的高频细节。本文提出ViSIR与ViFOR两种频域感知框架:ViSIR(视觉变压器-正弦隐式表示)结合视觉变压器与正弦激活函数以缓解频谱偏差;ViFOR(视觉变压器傅里叶表示网络)通过显式傅里叶滤波实现高低频独立学习。在包含地表温度、短波和长波通量的E3SM-HR地球系统数据集上评估显示,这两种模型优于领先的卷积神经网络、生成网络及原始Transformer基线,在峰值信噪比上最高提升2.6~dB,且结构相似性更高。
原文摘要 · Abstract (English)
Super-resolution can play an essential role in enhancing the spatial fidelity of Earth System Model outputs, allowing fine-scale structures highly beneficial to climate science to be recovered from coarse simulations. However, traditional deep super-resolution methods, including convolutional and transformer based models, tend to exhibit spectral bias, reconstructing low-frequency content more readily than valuable high-frequency details. In this work, we introduce ViSIR and ViFOR, two frequency-aware frameworks. ViSIR stands for the Vision Transformer-Tuned Sinusoidal Implicit Representation. ViSIR combines vision transformers with sinusoidal activations to mitigate spectral bias. ViFOR stands for the Vision Transformer Fourier Representation Network. ViFOR integrates explicit Fourier based filtering for independent low- and high-frequency learning. Evaluated on the E3SM-HR Earth system dataset across surface temperature, shortwave, and longwave fluxes, these models outperform leading Convolutional NN, Generative Networks, and vanilla transformer baselines, with ViFOR demonstrating up to 2.6~dB improvements in Peak Signal to Noise Ratio and higher Structural Similarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。