arXiv:2511.11706cs.LGcs.CV2025-11被引 1

融合多源遥感数据,实现10米分辨率高时频环境建模

Context-Aware Multimodal Representation Learning for Spatio-Temporally Explicit Environmental Modelling

  • 分两阶段融合哨兵1号与2号数据,保留传感器特异性并统一特征空间
  • 生成10米分辨率、云图间隔频率的潜空间,支持精细生态分析
  • 适合需要高时空精度的环境监测与生态系统动态研究

地球观测(EO)基础模型已成为从多种遥感传感器中提取地球系统潜在表示的有效方法。这些模型生成的嵌入可作为分析就绪数据集,无需大量传感器特定预处理即可建模生态系统动态。然而,现有模型通常在固定空间或时间尺度下运行,限制了对同时需要精细空间细节和高时间保真度的生态分析的应用。为克服此局限,我们提出一种表示学习框架,将不同EO模态整合至高时空分辨率的统一特征空间。以哨兵-1和哨兵-2数据为例,该方法在原生10米分辨率和无云哨兵-2采集频率下生成潜空间。各传感器先独立建模以捕捉其特异性,再融合至共享模型。这种两阶段设计支持模态特异性优化,并可轻松扩展至新传感器,仅重训练融合层而保留预训练编码器,从而捕捉互补遥感信息并保持时空一致性。定性分析显示,学习到的嵌入在异质景观中具有高空间与语义一致性。定量评估表明,其在估算总初级生产力方面编码了有意义的生态模式,且具备足够的时间保真度以支持细粒度分析。总体而言,该框架为需多样化时空分辨率的环境应用提供了灵活、分析就绪的表示学习方案。

原文摘要 · Abstract (English)

Earth observation (EO) foundation models have emerged as an effective approach to derive latent representations of the Earth system from various remote sensing sensors. These models produce embeddings that can be used as analysis-ready datasets, enabling the modelling of ecosystem dynamics without extensive sensor-specific preprocessing. However, existing models typically operate at fixed spatial or temporal scales, limiting their use for ecological analyses that require both fine spatial detail and high temporal fidelity. To overcome these limitations, we propose a representation learning framework that integrates different EO modalities into a unified feature space at high spatio-temporal resolution. We introduce the framework using Sentinel-1 and Sentinel-2 data as representative modalities. Our approach produces a latent space at native 10 m resolution and the temporal frequency of cloud-free Sentinel-2 acquisitions. Each sensor is first modeled independently to capture its sensor-specific characteristics. Their representations are then combined into a shared model. This two-stage design enables modality-specific optimisation and easy extension to new sensors, retaining pretrained encoders while retraining only fusion layers. This enables the model to capture complementary remote sensing data and to preserve coherence across space and time. Qualitative analyses reveal that the learned embeddings exhibit high spatial and semantic consistency across heterogeneous landscapes. Quantitative evaluation in modelling Gross Primary Production reveals that they encode ecologically meaningful patterns and retain sufficient temporal fidelity to support fine-scale analyses. Overall, the proposed framework provides a flexible, analysis-ready representation learning approach for environmental applications requiring diverse spatial and temporal resolutions.

遥感多模态环境建模时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。