融合遥感与雷达图像,提升土地覆盖分类精度与可解释性。
CLAIRE: A Dual Encoder Network with RIFT Loss and Phi-3 Small Language Model Based Interpretability for Cross-Modality Synthetic Aperture Radar and Optical Land Cover Segmentation
- 双编码器分别提取光学与雷达图像特征,通过跨模态注意力融合互补信息。
- 引入RIFT损失函数,在罕见类别上实现mIoU达86.86%、准确率94.58%。
- 基于Phi-3小模型生成预测理由,提升模型决策透明度,适合科研与应用落地。
从卫星影像中精确进行土地覆盖分类对环境监测和可持续资源管理至关重要。然而,自然景观复杂、类别视觉相似以及数据集中的显著类别不平衡,使该任务仍具挑战。为此,我们提出一种双编码器架构,独立提取光学与合成孔径雷达(SAR)图像的模态特异性特征,并通过名为CLAIRE的跨模态注意力融合模块进行融合。该机制突出互补的空间与纹理特征,使网络更有效地捕捉多样化的土地覆盖模式。我们采用混合损失函数,结合加权焦点损失与Tversky损失,命名为RIFT(Rare-Instance Focal-Tversky),以缓解类别不平衡并提升对少数类别的分割性能。模型在多个基准测试中表现优异:在WHU-OPT-SAR数据集上达到56.02%的mIoU与84.56%的总体准确率;在OpenEarthMap-SAR数据集上展现强泛化能力,mIoU为59.89%,准确率为73.92%;在云遮条件下表现卓越,PIE-RGB-SAR数据集上mIoU达86.86%,准确率94.58%。此外,我们引入由小语言模型(Phi-3)驱动的可解释性模块,生成专家级、样本相关的推理说明,显著增强模型决策透明性。
原文摘要 · Abstract (English)
Accurate land cover classification from satellite imagery is crucial in environmental monitoring and sustainable resource management. However, it remains challenging due to the complexity of natural landscapes, the visual similarity between classes, and the significant class imbalance in the available datasets. To address these issues, we propose a dual encoder architecture that independently extracts modality-specific features from optical and Synthetic Aperture Radar (SAR) imagery, which are then fused using a cross-modality attention-fusion module named Cross-modality Land cover segmentation with Attention and Imbalance-aware Reasoning-Enhanced Explanations (CLAIRE). This fusion mechanism highlights complementary spatial and textural features, enabling the network to better capture detailed and diverse land cover patterns. We incorporate a hybrid loss function that utilizes Weighted Focal Loss and Tversky Loss named RIFT (Rare-Instance Focal-Tversky) to address class imbalance and improve segmentation performance across underrepresented categories. Our model achieves competitive performance across multiple benchmarks: a mean Intersection over Union (mIoU) of 56.02% and Overall Accuracy (OA) of 84.56% on the WHU-OPT-SAR dataset; strong generalization with a mIoU of 59.89% and OA of 73.92% on the OpenEarthMap-SAR dataset; and remarkable robustness under cloud-obstructed conditions, achieving an mIoU of 86.86% and OA of 94.58% on the PIE-RGB-SAR dataset. Additionally, we introduce a metric-driven reasoning module generated by a Small Language Model (Phi-3), which generates expert-level, sample-specific justifications for model predictions, thereby enhancing transparency and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。