用Transformer从多光谱图生成高光谱图,兼顾清晰细节与真实光谱。
SpecSwin3D: Generating Hyperspectral Imagery from Multispectral Data via Transformer Networks
- 基于3D移位窗口Transformer,用5个波段输入重建224个波段的高光谱图像。
- 在相同分辨率下实现35.82dB PSNR、2.40° SAM,优于基线模型5.6dB以上。
- 适合农业监测、环境评估等需高光谱细节的应用场景。
多光谱和高光谱影像在农业、环境监测和城市规划中广泛应用,因其互补的空间与光谱特性。然而存在根本权衡:多光谱图像空间分辨率高但光谱信息少,而高光谱图像光谱丰富但空间分辨率低。以往的高光谱生成方法(如全色锐化变体、矩阵分解、卷积神经网络)常难以同时保持空间细节与光谱保真度。为此,本文提出SpecSwin3D,一种基于Transformer的模型,可从多光谱输入生成高光谱影像,同时保留空间与光谱质量。具体而言,SpecSwin3D以五个多光谱波段为输入,重建出224个高光谱波段且保持相同空间分辨率。此外,我们观察到远离输入波段的高光谱波段重建误差较大,因此引入级联训练策略逐步扩展光谱范围,稳定学习并提升保真度。同时,设计优化的波段序列,在3D移位窗口Transformer框架中重复并有序排列输入波段,以更好捕捉波段间关系。定量结果表明,本模型在PSNR达35.82 dB、SAM为2.40°、SSIM为0.96,相比基线模型MHF-Net提升5.6 dB PSNR,ERGAS降低超过一半。进一步在土地利用分类与烧毁区域分割两个下游任务上验证了其实际价值。
原文摘要 · Abstract (English)
Multispectral and hyperspectral imagery are widely used in agriculture, environmental monitoring, and urban planning due to their complementary spatial and spectral characteristics. A fundamental trade-off persists: multispectral imagery offers high spatial but limited spectral resolution, while hyperspectral imagery provides rich spectra at lower spatial resolution. Prior hyperspectral generation approaches (e.g., pan-sharpening variants, matrix factorization, CNNs) often struggle to jointly preserve spatial detail and spectral fidelity. In response, we propose SpecSwin3D, a transformer-based model that generates hyperspectral imagery from multispectral inputs while preserving both spatial and spectral quality. Specifically, SpecSwin3D takes five multispectral bands as input and reconstructs 224 hyperspectral bands at the same spatial resolution. In addition, we observe that reconstruction errors grow for hyperspectral bands spectrally distant from the input bands. To address this, we introduce a cascade training strategy that progressively expands the spectral range to stabilize learning and improve fidelity. Moreover, we design an optimized band sequence that strategically repeats and orders the five selected multispectral bands to better capture pairwise relations within a 3D shifted-window transformer framework. Quantitatively, our model achieves a PSNR of 35.82 dB, SAM of 2.40°, and SSIM of 0.96, outperforming the baseline MHF-Net by +5.6 dB in PSNR and reducing ERGAS by more than half. Beyond reconstruction, we further demonstrate the practical value of SpecSwin3D on two downstream tasks, including land use classification and burnt area segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。