统一处理遥感图像的时空谱差异,一模型适配多种任务与类别。
Spatial-Temporal-Spectral Unified Modeling for Remote Sensing Dense Prediction
- 通过可学习的元数据编码,统一建模任意尺寸、时长和波段的遥感数据。
- 单个模型在多个数据集上实现最优性能,支持多任务与可变类别预测。
- 适合需要灵活适配多源遥感数据的科研与应用开发者。
多源遥感数据的激增推动了密集预测深度学习的发展,但数据与任务统一仍面临挑战。现有遥感深度学习架构本质僵化,仅支持固定输入输出配置,难以适应真实数据中固有的异构时空谱特性。同时,这些模型忽视语义分割、二值变化检测与语义变化检测之间的内在关联,需为不同任务分别设计模型或解码器。此外,模型限定于预定义的语义类别集合,类别变更需昂贵的重新训练。为此,本文提出时空谱统一网络(STSUN),通过利用数据元信息实现任意空间尺寸、时间长度和光谱波段的统一表示。同时,通过可学习的任务嵌入,将语义分割、二值变化检测、语义变化检测等任务统一于单一架构;通过可学习的类别嵌入,实现对多组语义类别的灵活预测。在多种场景下的多个数据集上,实验表明单一STSUN模型能有效适应异构输入输出,统一多种密集预测任务与多样化语义类别,持续达到最先进性能,验证其在复杂遥感应用中的鲁棒性与泛化能力。
原文摘要 · Abstract (English)
The proliferation of multi-source remote sensing data has propelled the development of deep learning for dense prediction, yet significant challenges in data and task unification persist. Current deep learning architectures for remote sensing are fundamentally rigid. They are engineered for fixed input-output configurations, restricting their adaptability to the heterogeneous spatial, temporal, and spectral dimensions inherent in real-world data. Furthermore, these models neglect the intrinsic correlations among semantic segmentation, binary change detection, and semantic change detection, necessitating the development of distinct models or task-specific decoders. This paradigm is also constrained to a predefined set of output semantic classes, where any change to the classes requires costly retraining. To overcome these limitations, we introduce the Spatial-Temporal-Spectral Unified Network (STSUN) for unified modeling. STSUN can adapt to input and output data with arbitrary spatial sizes, temporal lengths, and spectral bands by leveraging their metadata for a unified representation. Moreover, STSUN unifies disparate dense prediction tasks within a single architecture by conditioning the model on trainable task embeddings. Similarly, STSUN facilitates flexible prediction across multiple set of semantic categories by integrating trainable category embeddings as metadata. Extensive experiments on multiple datasets with diverse Spatial-Temporal-Spectral configurations in multiple scenarios demonstrate that a single STSUN model effectively adapts to heterogeneous inputs and outputs, unifying various dense prediction tasks and diverse semantic class predictions. The proposed approach consistently achieves state-of-the-art performance, highlighting its robustness and generalizability for complex remote sensing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。