EarthMind融合多传感器遥感数据,提升环境监测理解能力。
EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
- 设计分层跨模态注意力机制,统一处理单与多传感器输入
- 在30,000对数据上训练,2,841对评测集上达领先性能
- 适合遥感分析、环境监测等需要多源数据融合的场景
地球观测(EO)数据分析对监测环境与人类活动至关重要。近年来的多模态大语言模型(MLLMs)虽在EO理解方面展现潜力,但仍局限于单一传感器输入,忽视异构模态间的互补性。我们提出EarthMind,一种统一的视觉-语言框架,通过创新的分层跨模态注意力(HCA)设计,可处理单传感器与跨传感器输入。HCA分层捕捉跨传感器的视觉关系,并将其与语言查询对齐,实现光学与合成孔径雷达(SAR)特征的自适应融合。为支持跨传感器学习,我们构建了包含3万对样本的FusionEO数据集,以及包含2,841对专家标注样本的EarthMind-Bench基准。大量实验表明,EarthMind在EarthMind-Bench上达到当前最优表现,并在多个EO基准上超越现有MLLMs。
原文摘要 · Abstract (English)
Earth Observation (EO) data analysis is vital for monitoring environmental and human dynamics. Recent Multimodal Large Language Models (MLLMs) show potential in EO understanding but remain restricted to single-sensor inputs, overlooking the complementarity across heterogeneous modalities. We propose EarthMind, a unified vision-language framework that handles both single- and cross-sensor inputs via an innovative hierarchical cross-modal attention (ie, HCA) design. Specifically, HCA hierarchically captures visual relationships across sensors and aligns them with language queries, enabling adaptive fusion of optical and Synthetic Aperture Radar (SAR) features. To support cross-sensor learning, we curate FusionEO, a 30K-pair dataset with diverse annotations, and establish EarthMind-Bench, a 2,841-pair benchmark with expert annotations for perception and reasoning tasks. Extensive experiments show that EarthMind achieves state-of-the-art results on EarthMind-Bench and surpasses existing MLLMs on multiple EO benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。