arXiv:2605.00296cs.CV2026-05

用视觉变压器高效分类植被时序像素,省算力还抗长序列

Efficient Spatio-Temporal Vegetation Pixel Classification with Vision Transformers

论文配图:Efficient Spatio-Temporal Vegetation Pixel Classification with Vision Transformers
图 1 · 摘自论文原文
  • 改用视觉变压器替代传统卷积网络,优化七项设计提升效率
  • 计算量降低一个数量级,参数量不随时间序列变长而增加
  • 适合资源受限的长期生态监测系统,尤其适用于无人机数据

植物物候学——研究周期性生命事件——对理解生态系统动态及其对气候变化的响应至关重要。尽管无人机和近地面相机可实现高分辨率监测,但跨时间识别植物物种仍面临计算挑战。现有先进方法如多时相卷积网络依赖僵化的多分支结构,随时间序列延长而扩展性差,且需大空间上下文窗口。本文针对视觉变压器(ViT)在高效时空植被像素分类中的应用展开深入研究,系统分析了七项关键设计维度:(i) 数据归一化;(ii) 光谱排列;(iii) 边界处理;(iv) 空间上下文窗口形状与大小;(v) 分块策略;(vi) 位置编码;(vii) 特征聚合方式。方法在巴西塞拉多生物群落的两个数据集上评估:Serra do Cipó(航拍影像)和Itirapina(近地面影像)。实验表明,所提ViT方法在保持竞争力分类性能的同时,显著提升计算效率。特别地,该方法将浮点运算量(FLOPs)降低一个数量级,且参数量不随时间序列长度变化,而卷积网络基线则呈线性增长。结果证实,视觉变压器是资源受限物候监测系统的稳健、可扩展解决方案。

原文摘要 · Abstract (English)

Plant phenology-the study of recurrent life cycle events-is essential for understanding ecosystem dynamics and their responses to climate change impacts. While Unmanned Aerial Vehicles (UAVs) and near-surface cameras enable high-resolution monitoring, identifying plant species across time remains computationally challenging. State-of-the-art approaches, specifically Multi-Temporal Convolutional Networks (CNNs), rely on rigid multi-branch architectures that scale poorly with longer time series and require large spatial context windows. In this paper, we present an extensive study on optimizing Vision Transformers (ViTs) for efficient spatio-temporal vegetation pixel classification. We conducted a comprehensive ablation study analyzing seven key design dimensions, including: (i) data normalization; (ii) spectral arrangement; (iii) boundary handling; (iv) spatial context window shape and size; (v) tokenization strategies; (vi) positional encoding; and (vii) feature aggregation strategies. Our method was evaluated on two datasets from the Brazilian Cerrado biome, Serra do Cipó (aerial imagery) and Itirapina (near-surface imagery). Experimental results demonstrate that our ViT approach offers a substantial improvement in computational efficiency while maintaining competitive classification performance. Notably, our ViT reduces Floating Point Operations (FLOPs) by an order of magnitude and maintains constant parameter complexity regardless of the time series length, whereas the CNN baseline scales linearly. Our findings confirm that ViTs are a robust, scalable solution for resource-constrained phenological monitoring systems.

视觉变压器物候监测高效计算遥感分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。