用空间位置信息融合遥感、物种记录等异构生态数据,提升环境预测精度。
A Geolocation-Aware Multimodal Approach for Ecological Prediction
- 基于Transformer的多模态融合,用位置感知嵌入保留样本间空间关系。
- 在瑞士103个环境变量预测中,多模态比单模态平均提升8.7%性能。
- 适合做生态监测与大范围环境制图的研究者使用。
虽然融合多种模态有望提升环境监测能力,但现有方法难以处理格式或内容各异的数据。核心挑战在于如何结合连续网格数据(如遥感影像)与稀疏不规则的点状观测(如物种记录)。当前地理统计与深度学习方法通常仅处理单一模态或空间对齐输入,无法有效解决此问题。本文提出一种位置感知多模态方法(GAMMA),基于Transformer架构,通过显式空间上下文整合异构生态数据。GAMMA不将观测插值到统一网格,而是先将所有输入表示为保留样本间空间关系的位置感知嵌入,并动态跨模态与空间尺度选择相关邻域,实现遥感影像与稀疏地理标记观测的联合利用。我们在瑞士的SWECO25数据立方体上评估了该方法,任务是预测103个环境变量,输入包括航空影像、来自GBIF的生物多样性观测以及EcoWikiRS数据集提供的维基百科文本生境描述。实验表明,多模态融合持续优于单模态基线,且显式空间上下文进一步提升了模型准确率。GAMMA灵活的架构还支持通过消融实验分析各模态贡献。结果证明,位置感知多模态学习在整合异构生态数据及支持大规模环境制图与生物多样性监测方面具有潜力。
原文摘要 · Abstract (English)
While integrating multiple modalities has the potential to improve environmental monitoring, current approaches struggle to combine data sources with heterogeneous formats or contents. A central difficulty arises when combining continuous gridded data (e.g., remote sensing) with sparse and irregular point observations such as species records. Existing geostatistical and deep-learning-based approaches typically operate on a single modality or focus on spatially aligned inputs, and thus cannot seamlessly overcome this difficulty. We propose a Geolocation-Aware MultiModal Approach (GAMMA), a transformer-based fusion approach designed to integrate heterogeneous ecological data using explicit spatial context. Instead of interpolating observations into a common grid, GAMMA first represents all inputs as location-aware embeddings that preserve spatial relationships between samples. GAMMA dynamically selects relevant neighbours across modalities and spatial scales, enabling the model to jointly exploit continuous remote sensing imagery and sparse geolocated observations. We evaluate GAMMA on the task of predicting 103 environmental variables from the SWECO25 data cube across Switzerland. Inputs combine aerial imagery with biodiversity observations from GBIF and textual habitat descriptions from Wikipedia, provided by the EcoWikiRS dataset. Experiments show that multimodal fusion consistently improves prediction performance over single-modality baselines and that explicit spatial context further enhances model accuracy. The flexible architecture of GAMMA also allows to analyse the contribution of each modality through controlled ablation experiments. These results demonstrate the potential of location-aware multimodal learning for integrating heterogeneous ecological data and for supporting large-scale environmental mapping tasks and biodiversity monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。