arXiv:2503.08534cs.CVcs.LG2025-03

提出多光谱注意力机制,大幅提升遥感地表分类精度。

ChromaFormer: A Scalable and Accurate Transformer Architecture for Land Cover Classification

  • 设计新型多光谱注意力机制,适配十通道以上遥感数据
  • 6.55亿参数模型在比利时全境数据上达95%以上准确率
  • 适合大规模高分辨率遥感图像分析任务

Sentinel等遥感系统可提供约10米分辨率的全球覆盖影像。深度学习模型在UCMerced和ISPRS Vaihingen等基准上表现优异,但传统卷积模型如UNet、ResNet仅支持三通道RGB数据,难以利用卫星提供的十余个波段。尽管已有若干变压器架构用于遥感,但缺乏广泛验证,且多在小规模数据集(如Salinas Valley)上测试。随着一国一级行政区划的密集空间土地利用标签成为可能,规模化模型的潜力凸显。本文提出ChromaFormer,一类多光谱变压器模型,在覆盖超过13,500 km²、含15类的地表分类数据集——比利时弗兰德斯生物估值图上进行跨数量级参数规模评估。提出新颖的多光谱注意力策略并通过消融实验证明其有效性。结果表明,远超传统架构的大型模型显著提升性能:2300万参数的UNet++模型准确率不足65%,而6.55亿参数的多光谱变压器达到95%以上准确率。

原文摘要 · Abstract (English)

Remote sensing imagery from systems such as Sentinel provides full coverage of the Earth's surface at around 10-meter resolution. The remote sensing community has transitioned to extensive use of deep learning models due to their high performance on benchmarks such as the UCMerced and ISPRS Vaihingen datasets. Convolutional models such as UNet and ResNet variations are commonly employed for remote sensing but typically only accept three channels, as they were developed for RGB imagery, while satellite systems provide more than ten. Recently, several transformer architectures have been proposed for remote sensing, but they have not been extensively benchmarked and are typically used on small datasets such as Salinas Valley. Meanwhile, it is becoming feasible to obtain dense spatial land-use labels for entire first-level administrative divisions of some countries. Scaling law observations suggest that substantially larger multi-spectral transformer models could provide a significant leap in remote sensing performance in these settings. In this work, we propose ChromaFormer, a family of multi-spectral transformer models, which we evaluate across orders of magnitude differences in model parameters to assess their performance and scaling effectiveness on a densely labeled imagery dataset of Flanders, Belgium, covering more than 13,500 km^2 and containing 15 classes. We propose a novel multi-spectral attention strategy and demonstrate its effectiveness through ablations. Furthermore, we show that models many orders of magnitude larger than conventional architectures, such as UNet, lead to substantial accuracy improvements: a UNet++ model with 23M parameters achieves less than 65% accuracy, while a multi-spectral transformer with 655M parameters achieves over 95% accuracy on the Biological Valuation Map of Flanders.

遥感分类多光谱变压器大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。