提出联合重建与预测框架,提升云遮挡下遥感影像分割性能
A Joint Learning Framework with Feature Reconstruction and Prediction for Incomplete Satellite Image Time Series in Agricultural Semantic Segmentation
- 用时序掩码模拟缺失,联合训练特征重建与分类任务
- 在湖南、法国等地实验中,作物提取F1提升6.93%,分类提升7.09%
- 可适配不同传感器和缺失率,避免冗余重建与噪声传播
卫星影像时序序列(SITS)对农业语义分割至关重要。然而,云污染导致时间片段缺失,破坏时序依赖关系并引发特征偏移,使基于完整SITS训练的模型性能下降。现有方法通常先完整重建SITS再进行预测,或通过数据增强模拟缺失数据,但全量重建易引入噪声与冗余,增强模型仅能处理有限缺失模式,泛化能力差。本文提出一种联合学习框架,同时进行特征重建与预测。训练时使用时序掩码模拟缺失场景,两个任务由真实标签和在完整SITS上训练的教师模型共同指导。预测任务约束模型从掩码输入中选择性重建与教师时序特征一致的关键特征,减少不必要的重建并控制噪声传播。通过将重建特征融入预测任务,模型避免学习捷径,保持对多样化缺失模式及完整数据的处理能力。在湖南、西法兰西和加泰罗尼亚地区的SITS数据上实验表明,该方法在作物提取任务中平均F1得分提升6.93%,作物分类提升7.09%。模型在不同卫星传感器(含Sentinel-2、PlanetScope)下,于不同时间缺失率和模型主干结构下均表现良好。
原文摘要 · Abstract (English)
Satellite Image Time Series (SITS) is crucial for agricultural semantic segmentation. However, Cloud contamination introduces time gaps in SITS, disrupting temporal dependencies and causing feature shifts, leading to degraded performance of models trained on complete SITS. Existing methods typically address this by reconstructing the entire SITS before prediction or using data augmentation to simulate missing data. Yet, full reconstruction may introduce noise and redundancy, while the data-augmented model can only handle limited missing patterns, leading to poor generalization. We propose a joint learning framework with feature reconstruction and prediction to address incomplete SITS more effectively. During training, we simulate data-missing scenarios using temporal masks. The two tasks are guided by both ground-truth labels and the teacher model trained on complete SITS. The prediction task constrains the model from selectively reconstructing critical features from masked inputs that align with the teacher's temporal feature representations. It reduces unnecessary reconstruction and limits noise propagation. By integrating reconstructed features into the prediction task, the model avoids learning shortcuts and maintains its ability to handle varied missing patterns and complete SITS. Experiments on SITS from Hunan Province, Western France, and Catalonia show that our method improves mean F1-scores by 6.93% in cropland extraction and 7.09% in crop classification over baselines. It also generalizes well across satellite sensors, including Sentinel-2 and PlanetScope, under varying temporal missing rates and model backbones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。