EarthDial让遥感数据能用自然语言对话方式分析,支持多光谱、多时相、多分辨率图像。
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
- 构建可对话的遥感分析系统,融合RGB、SAR、近红外等多模态数据
- 在44个下游任务中超越现有模型,实现跨任务更好泛化性能
- 适合环境监测、灾害响应和资源管理领域的研究人员与应用者
通过交互式视觉-语言模型(VLMs)自动分析海量地球观测数据,可为环境监测、灾害响应和资源管理带来新机遇。现有通用VLMs在遥感数据上表现不佳,而近期地理空间VLMs仍受限于固定分辨率和有限传感器模态。本文提出EarthDial,一个专为地球观测(EO)数据设计的对话助手,将复杂的多感官地球观测转化为自然语言交互对话。EarthDial支持多光谱、多时相和多分辨率影像,涵盖分类、检测、描述生成、问答、视觉推理和视觉定位等多种遥感任务。为此,我们构建了包含超过1111万条指令对的大型指令微调数据集,覆盖RGB、合成孔径雷达(SAR)以及近红外(NIR)和红外等多模态数据。此外,EarthDial可处理双时相与多时相序列分析,适用于变化检测等应用。在44个下游数据集上的广泛实验表明,EarthDial在各类遥感任务中均优于现有通用及领域专用模型,展现出更优的泛化能力。源代码与预训练模型见https://github.com/hiyamdebary/EarthDial。
原文摘要 · Abstract (English)
Automated analysis of vast Earth observation data via interactive Vision-Language Models (VLMs) can unlock new opportunities for environmental monitoring, disaster response, and {resource management}. Existing generic VLMs do not perform well on Remote Sensing data, while the recent Geo-spatial VLMs remain restricted to a fixed resolution and few sensor modalities. In this paper, we introduce EarthDial, a conversational assistant specifically designed for Earth Observation (EO) data, transforming complex, multi-sensory Earth observations into interactive, natural language dialogues. EarthDial supports multi-spectral, multi-temporal, and multi-resolution imagery, enabling a wide range of remote sensing tasks, including classification, detection, captioning, question answering, visual reasoning, and visual grounding. To achieve this, we introduce an extensive instruction tuning dataset comprising over 11.11M instruction pairs covering RGB, Synthetic Aperture Radar (SAR), and multispectral modalities such as Near-Infrared (NIR) and infrared. Furthermore, EarthDial handles bi-temporal and multi-temporal sequence analysis for applications like change detection. Our extensive experimental results on 44 downstream datasets demonstrate that EarthDial outperforms existing generic and domain-specific models, achieving better generalization across various EO tasks. Our source codes and pre-trained models are at https://github.com/hiyamdebary/EarthDial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。