用高分辨率影像与高程图联合预训练,提升建筑级遥感分析性能。
HiRes-FusedMIM: A High-Resolution RGB-DSM Pre-trained Model for Building-Level Remote Sensing Applications
- 双编码器结构分别处理可见光与高程数据,融合多目标损失学习联合表征。
- 在WHU、LoveDA等数据集上超越现有方法,建筑细节识别能力显著提升。
- 高程信息对建筑分析至关重要,适合数字孪生、城市建模等应用者。
自监督学习的发展催生了众多基础模型,在计算机视觉任务中表现卓越。然而,这些模型常忽视高分辨率数字表面模型(DSM)在城市环境理解中的关键作用,尤其在建筑级分析中,这对数字孪生等应用至关重要。为此,我们提出HiRes-FusedMIM,一种专为高分辨率RGB与DSM数据设计的预训练模型。该模型采用双编码器简单掩码图像建模(SimMIM)架构,结合重建与对比学习的多目标损失函数,实现两模态信息的深层融合。我们在分类、语义分割和实例分割等下游任务上进行评估,结果表明:1)在WHU航空影像与LoveDA等建筑相关数据集上,该模型性能优于当前最优地理空间方法,证明其能有效捕捉细粒度建筑信息;2)预训练阶段引入DSM可持续提升性能,凸显高程信息对建筑分析的价值;3)双编码器结构在Vaihingen分割任务上显著优于单编码器模型,说明各模态专用表征学习的优势。为促进后续研究,模型权重将公开发布。
原文摘要 · Abstract (English)
Recent advances in self-supervised learning have led to the development of foundation models that have significantly advanced performance in various computer vision tasks. However, despite their potential, these models often overlook the crucial role of high-resolution digital surface models (DSMs) in understanding urban environments, particularly for building-level analysis, which is essential for applications like digital twins. To address this gap, we introduce HiRes-FusedMIM, a novel pre-trained model specifically designed to leverage the rich information contained within high-resolution RGB and DSM data. HiRes-FusedMIM utilizes a dual-encoder simple masked image modeling (SimMIM) architecture with a multi-objective loss function that combines reconstruction and contrastive objectives, enabling it to learn powerful, joint representations from both modalities. We conducted a comprehensive evaluation of HiRes-FusedMIM on a diverse set of downstream tasks, including classification, semantic segmentation, and instance segmentation. Our results demonstrate that: 1) HiRes-FusedMIM outperforms previous state-of-the-art geospatial methods on several building-related datasets, including WHU Aerial and LoveDA, demonstrating its effectiveness in capturing and leveraging fine-grained building information; 2) Incorporating DSMs during pre-training consistently improves performance compared to using RGB data alone, highlighting the value of elevation information for building-level analysis; 3) The dual-encoder architecture of HiRes-FusedMIM, with separate encoders for RGB and DSM data, significantly outperforms a single-encoder model on the Vaihingen segmentation task, indicating the benefits of learning specialized representations for each modality. To facilitate further research and applications in this direction, we will publicly release the trained model weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。