arXiv:2412.16583cs.CV2024-12被引 11

让视觉语言模型学会精准估算地球观测中的生物量。

REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation

  • 用语言引导推理,结合科学知识提升模型对地理数据的回归能力
  • 构建160万组多模态遥感图像与文本对,支持生物量回归与内容生成
  • 适合环境监测、资源管理领域的研究人员和工程师使用

视觉语言模型(VLM)的快速发展推动了人工智能在地球观测(EO)领域的进步,但其应用仍主要集中在图像内容描述。本文提出全新的基准数据集REO-Instruct,包含160万组多模态地球观测图像与语言配对,统一支持生物量回归与图像内容解释任务。基于该数据集,我们开发了REO-VLM,首次将回归能力与传统生成功能无缝融合。通过语言驱动的推理机制引入科学领域知识,模型不再仅依赖遥感图像,实现对复杂科学属性的综合解析,显著提升环境监测与资源管理能力,树立新性能标杆。

原文摘要 · Abstract (English)

The rapid evolution of Vision Language Models (VLMs) has catalyzed significant advancements in artificial intelligence, expanding research across various disciplines, including Earth Observation (EO). While VLMs have enhanced image understanding and data processing within EO, their applications have predominantly focused on image content description. This limited focus overlooks their potential in geographic and scientific regression tasks, which are essential for diverse EO applications. To bridge this gap, this paper introduces a novel benchmark dataset, called \textbf{REO-Instruct} to unify regression and generation tasks specifically for the EO domain. Comprising 1.6 million multimodal EO imagery and language pairs, this dataset is designed to support both biomass regression and image content interpretation tasks. Leveraging this dataset, we develop \textbf{REO-VLM}, a groundbreaking model that seamlessly integrates regression capabilities with traditional generative functions. By utilizing language-driven reasoning to incorporate scientific domain knowledge, REO-VLM goes beyond solely relying on EO imagery, enabling comprehensive interpretation of complex scientific attributes from EO data. This approach establishes new performance benchmarks and significantly enhances the capabilities of environmental monitoring and resource management.

地球观测视觉语言模型回归任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。