arXiv:2412.01944cs.CVeess.IV2024-12被引 4

对比卷积与注意力模型在卫星时序图像作物分割中的表现

A Comparative Study of Transformer and Convolutional Models for Crop Segmentation from Satellite Image Time Series

  • 用3D CNN和三种变压器架构比较时序遥感图像分割方法
  • TSViT在慕尼黑和伦巴第数据集上表现最佳,优于3D U-Net
  • VistaFormer效率最高,适合资源受限场景

从卫星图像时序序列(SITS)中进行作物分割是农业监测与土地利用分析的基础任务。尽管卷积神经网络(CNN)已广泛应用,但基于Transformer的架构为多光谱数据中的时空依赖建模提供了新路径。本文对比了3D U-Net、3D FPN、3D DeepLabv3等3D CNN模型,以及Swin UNETR、TSViT、VistaFormer三种基于Transformer的模型在哨兵2号时序数据上的作物分类表现。在慕尼黑与伦巴第数据集上的实验表明,TSViT整体表现最优,略胜于3D U-Net这一强基线;VistaFormer在效率方面表现最佳,而Swin UNETR虽具竞争力,但不如显式建模时间动态的模型。结果表明,对时间维度的有效建模至关重要:TSViT超越了将时间视为额外空间维度的方法,接近最优性能,而VistaFormer在性能与效率间取得良好平衡。

原文摘要 · Abstract (English)

Crop segmentation from satellite image time series (SITS) is a fundamental task for agricultural monitoring and land-use analysis. While convolutional neural networks (CNNs) have been widely used, transformer-based architectures offer alternative mechanisms for representing spatial and temporal dependencies in multispectral data. This paper presents a comparative study of CNN and transformer-based segmentation models for crop mapping from Sentinel-2 time series, including 3D U-Net, 3D FPN, 3D DeepLabv3, and three transformer architectures: Swin UNETR, TSViT, and VistaFormer, which adopt different strategies for capturing temporal dependencies. Experiments on the Munich and Lombardia datasets show that TSViT achieves the best overall results, slightly surpassing 3D U-Net, which remains a strong CNN baseline. VistaFormer offers the best efficiency, while Swin UNETR performs competitively but is less effective than transformers that explicitly model temporal dynamics. These results highlight that temporal modelling is critical for SITS: TSViT outperforms CNNs and approaches that treat time as an additional spatial dimension, while VistaFormer provides a strong efficiency-performance trade-off.

作物分割时序图像变压器遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。