解决遥感影像时间错配导致的分割误差,提升高分辨率树冠识别准确率。
Multi-modal Uncertainty Robust Tree Cover Segmentation For High-Resolution Remote Sensing Images
- 用主模态重建不确定的辅模态特征,增强多源数据融合鲁棒性。
- 在上海与苏黎世数据集上,整体分割精度提升6.2%以上。
- 适合需要高精度树冠制图的生态监测与城市规划场景。
多模态遥感图像语义分割的进展显著提升了树冠制图精度,支持城市规划、森林监测与生态评估。融合光学影像、激光雷达(LiDAR)和合成孔径雷达(SAR)等多源数据相比单模态方法表现更优。然而,这些数据常因采集时间间隔数日甚至数月,导致植被扰动(如砍伐、火灾)或成像质量变化,引入跨模态不确定性,尤其在高分辨率图像中严重降低分割性能。为此,本文提出MURTreeFormer,一种新型多模态分割框架,通过建模辅模态的局部不确定性并加以利用,实现鲁棒树冠分割。该框架将一模态设为主模态,其余为辅模态,通过概率潜在表示显式建模辅模态的块级不确定性;对不确定块,基于变分自编码器(VAE)重采样机制,从主模态分布中重建特征,生成增强的辅模态特征用于融合。解码器中引入梯度幅值注意力(GMA)模块与轻量级精修头(RH),引导关注树状结构并保留细粒度空间细节。在上海与苏黎世多模态数据集上的大量实验表明,MURTreeFormer显著提升分割性能,并有效缓解时序引起的异质性不确定性。
原文摘要 · Abstract (English)
Recent advances in semantic segmentation of multi-modal remote sensing images have significantly improved the accuracy of tree cover mapping, supporting applications in urban planning, forest monitoring, and ecological assessment. Integrating data from multiple modalities-such as optical imagery, light detection and ranging (LiDAR), and synthetic aperture radar (SAR)-has shown superior performance over single-modality methods. However, these data are often acquired days or even months apart, during which various changes may occur, such as vegetation disturbances (e.g., logging, and wildfires) and variations in imaging quality. Such temporal misalignments introduce cross-modal uncertainty, especially in high-resolution imagery, which can severely degrade segmentation accuracy. To address this challenge, we propose MURTreeFormer, a novel multi-modal segmentation framework that mitigates and leverages aleatoric uncertainty for robust tree cover mapping. MURTreeFormer treats one modality as primary and others as auxiliary, explicitly modeling patch-level uncertainty in the auxiliary modalities via a probabilistic latent representation. Uncertain patches are identified and reconstructed from the primary modality's distribution through a VAE-based resampling mechanism, producing enhanced auxiliary features for fusion. In the decoder, a gradient magnitude attention (GMA) module and a lightweight refinement head (RH) are further integrated to guide attention toward tree-like structures and to preserve fine-grained spatial details. Extensive experiments on multi-modal datasets from Shanghai and Zurich demonstrate that MURTreeFormer significantly improves segmentation performance and effectively reduces the impact of temporally induced aleatoric uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。