用时空管状嵌入提升云遮盖下多光谱影像重建质量
Temporal-Spatial Tubelet Embedding for Cloud-Robust MSI Reconstruction using MSI-SAR Fusion: A Multi-Head Self-Attention Video Vision Transformer Approach
- 采用3D卷积提取短时序管状块,保持局部时间一致性
- 融合雷达与多光谱数据,使均方误差降低10.33%
- 适合农业遥感中云干扰严重区域的影像修复
多光谱影像(MSI)中的云覆盖严重影响早季作物制图,因破坏了光谱信息。现有基于视觉变压器(ViT)的时间序列重建方法,如SMTS-ViT,常使用粗粒度的时间嵌入,对整个序列进行聚合,导致信息大量丢失,降低重建精度。本研究提出一种基于视频视觉变压器(ViViT)的框架,结合时空融合嵌入,用于云覆盖区域的多光谱影像重建。通过3D卷积以受限时间跨度(t=2)提取非重叠管状块,确保局部时间连贯性,同时减少跨日信息退化。实验涵盖仅多光谱(MSI-only)与雷达-多光谱融合(SAR-MSI)两种场景。在2020年特拉尔县数据上的综合实验表明:MTS-ViViT相比MTS-ViT基线均方误差降低2.23%;而引入雷达数据后,SMTS-ViViT相较SMTS-ViT基线性能提升10.33%。该框架有效提升了云干扰环境下光谱重建质量,增强农业监测鲁棒性。
原文摘要 · Abstract (English)
Cloud cover in multispectral imagery (MSI) significantly hinders early-season crop mapping by corrupting spectral information. Existing Vision Transformer(ViT)-based time-series reconstruction methods, like SMTS-ViT, often employ coarse temporal embeddings that aggregate entire sequences, causing substantial information loss and reducing reconstruction accuracy. To address these limitations, a Video Vision Transformer (ViViT)-based framework with temporal-spatial fusion embedding for MSI reconstruction in cloud-covered regions is proposed in this study. Non-overlapping tubelets are extracted via 3D convolution with constrained temporal span $(t=2)$, ensuring local temporal coherence while reducing cross-day information degradation. Both MSI-only and SAR-MSI fusion scenarios are considered during the experiments. Comprehensive experiments on 2020 Traill County data demonstrate notable performance improvements: MTS-ViViT achieves a 2.23\% reduction in MSE compared to the MTS-ViT baseline, while SMTS-ViViT achieves a 10.33\% improvement with SAR integration over the SMTS-ViT baseline. The proposed framework effectively enhances spectral reconstruction quality for robust agricultural monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。