首个融合描述与回归的遥感生态分析基准,推动视觉语言模型科学推理
Towards Unified Vision Language Models for Forest Ecological Analysis in Earth Observation
- 构建统一遥感生态数据集,支持从分类到生物量估算的多任务评测
- 现有模型在数值推理上表现差,暴露出科学级视觉语言模型的短板
- 适合研究地球观测、生态建模与多模态模型泛化能力的学者
视觉语言模型在感知与推理方面取得显著进展,但在地球观测(EO)领域的科学回归任务中潜力尚未被充分挖掘。现有EO数据集主要聚焦于图像描述或分类等语义理解任务,缺乏将多模态感知与可度量生物物理变量对齐的基准。为此,我们提出首个统一基准REO-Instruct,同时支持描述性与回归性任务。该数据集在森林生态场景中建立可解释的逻辑链,涵盖人类活动识别、地表覆盖分类、生态斑块计数及地上生物量(AGB)回归。数据集整合了共配准的哨兵-2与ALOS-2影像,并通过人机协作管道生成与验证结构化文本标注。对通用视觉语言模型的全面评估显示,当前模型在数值推理方面表现不佳,凸显科学型视觉语言模型的核心挑战。REO-Instruct为下一代地理空间模型的研发与评估提供标准化基础。项目页面公开:https://github.com/zhu-xlab/REO-Instruct
原文摘要 · Abstract (English)
Recent progress in vision language models (VLMs) has enabled remarkable perception and reasoning capabilities, yet their potential for scientific regression in Earth Observation (EO) remains largely unexplored. Existing EO datasets mainly emphasize semantic understanding tasks such as captioning or classification, lacking benchmarks that align multimodal perception with measurable biophysical variables. To fill this gap, we present REO-Instruct, the first unified benchmark designed for both descriptive and regression tasks in EO. REO-Instruct establishes a cognitively interpretable logic chain in forest ecological scenario (human activity,land-cover classification, ecological patch counting, above-ground biomass (AGB) regression), bridging qualitative understanding and quantitative prediction. The dataset integrates co-registered Sentinel-2 and ALOS-2 imagery with structured textual annotations generated and validated through a hybrid human AI pipeline. Comprehensive evaluation protocols and baseline results across generic VLMs reveal that current models struggle with numeric reasoning, highlighting an essential challenge for scientific VLMs. REO-Instruct offers a standardized foundation for developing and assessing next-generation geospatial models capable of both description and scientific inference. The project page are publicly available at \href{https://github.com/zhu-xlab/REO-Instruct}{REO-Instruct}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。