arXiv:2604.19591cs.CV2026-04被引 2

解耦空间结构与语义,提升高分辨率遥感地图的全局一致性

Structure-Semantic Decoupled Modulation of Global Geospatial Embeddings for High-Resolution Remote Sensing Mapping

论文配图:Structure-Semantic Decoupled Modulation of Global Geospatial Embeddings for High-Resolution Remote Sensing Mapping
图 1 · 摘自论文原文
  • 将全局地理表征拆分为结构先验和语义注入两条路径
  • 在多种场景下显著提升复杂地物分类精度,减少碎片化预测
  • 适合需要融合大范围地理信息的高分辨率遥感应用

细粒度高分辨率遥感制图通常依赖局部视觉特征,限制了跨域泛化能力,并常导致大规模地物覆盖的碎片化预测。尽管全局地理基础模型具备强大的通用表征能力,但直接将其高维隐式嵌入与高分辨率视觉特征融合,常因严重的语义-空间鸿沟引发特征干扰和空间结构退化。为此,我们提出结构-语义解耦调制(SSDM)框架,将全局地理表征解耦为两条互补的跨模态注入路径。首先,结构先验调制分支将全局表示中的宏观感受野先验引入高分辨率编码器的自注意力模块,通过整体结构约束引导局部特征提取,有效抑制高频细节噪声和类内过强差异引起的预测碎片化。其次,全局语义注入分支显式对齐整体上下文与深层高分辨率特征空间,通过跨模态融合直接补充全局语义,显著增强复杂地物的语义一致性和类别区分能力。大量实验表明,本方法在跨模态融合任务中达到当前最优性能。通过释放全局嵌入潜力,SSDM在多种场景下持续提升高分辨率制图精度,为地理基础模型融入高分辨率视觉任务提供了通用且高效的范式。

原文摘要 · Abstract (English)

Fine-grained high-resolution remote sensing mapping typically relies on localized visual features, which restricts cross-domain generalizability and often leads to fragmented predictions of large-scale land covers. While global geospatial foundation models offer powerful, generalizable representations, directly fusing their high-dimensional implicit embeddings with high-resolution visual features frequently triggers feature interference and spatial structure degradation due to a severe semantic-spatial gap. To overcome these limitations, we propose a Structure-Semantic Decoupled Modulation (SSDM) framework, which decouples global geospatial representations into two complementary cross-modal injection pathways. First, the structural prior modulation branch introduces the macroscopic receptive field priors from global representations into the self-attention modules of the high-resolution encoder. By guiding local feature extraction with holistic structural constraints, it effectively suppresses prediction fragmentation caused by high-frequency detail noise and excessive intra-class variance. Second, the global semantic injection branch explicitly aligns holistic context with the deep high-resolution feature space and directly supplements global semantics via cross-modal integration, thereby significantly enhancing the semantic consistency and category-level discrimination of complex land covers. Extensive experiments demonstrate that our method achieves state-of-the-art performance compared to existing cross-modal fusion approaches. By unleashing the potential of global embeddings, SSDM consistently improves high-resolution mapping accuracy across diverse scenarios, providing a universal and effective paradigm for integrating geospatial foundation models into high-resolution vision tasks.

遥感制图跨模态融合地理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。