用像素级时序图提升遥感图像特征提取效果
Pixel-Wise Multimodal Contrastive Learning for Remote Sensing Images
- 将植被指数时序转为二维递归图,增强像素级表征
- 对比学习使遥感图像与时序特征表示质量显著提升
- 适合遥感变化检测、土地覆盖分类等任务
卫星持续生成海量地球观测数据,尤其是卫星图像时间序列(SITS)。现有深度学习模型多处理整幅图像或完整时序,难以捕捉像素级动态变化。本文提出一种新型像素级多模态对比学习方法——PIMC,通过将像素级植被指数时序(NDVI、EVI、SAVI)转换为递归图,生成更具信息量的二维表示,并结合遥感图像(RSI)进行自监督训练。在PASTIS数据集上评估了像素级预测与分类,在EuroSAT上评估土地覆盖分类任务。实验表明,二维表示显著提升特征表达能力,对比学习进一步优化了时序与图像的表示质量。该方法在多个地球观测任务中超越现有先进模型,为处理SITS与RSI提供了稳健的自监督框架。
原文摘要 · Abstract (English)
Satellites continuously generate massive volumes of data, particularly for Earth observation, including satellite image time series (SITS). However, most deep learning models are designed to process either entire images or complete time series sequences to extract meaningful features for downstream tasks. In this study, we propose a novel multimodal approach that leverages pixel-wise two-dimensional (2D) representations to encode visual property variations from SITS more effectively. Specifically, we generate recurrence plots from pixel-based vegetation index time series (NDVI, EVI, and SAVI) as an alternative to using raw pixel values, creating more informative representations. Additionally, we introduce PIxel-wise Multimodal Contrastive (PIMC), a new multimodal self-supervision approach that produces effective encoders based on two-dimensional pixel time series representations and remote sensing imagery (RSI). To validate our approach, we assess its performance on three downstream tasks: pixel-level forecasting and classification using the PASTIS dataset, and land cover classification on the EuroSAT dataset. Moreover, we compare our results to state-of-the-art (SOTA) methods on all downstream tasks. Our experimental results show that the use of 2D representations significantly enhances feature extraction from SITS, while contrastive learning improves the quality of representations for both pixel time series and RSI. These findings suggest that our multimodal method outperforms existing models in various Earth observation tasks, establishing it as a robust self-supervision framework for processing both SITS and RSI. Code avaliable on
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。