构建实时灾害智能评估基准,测试大模型从原始遥感数据中快速推理灾害的能力。
Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams

- 直接使用原始高频卫星数据与地面观测,跳过延迟处理流程。
- 覆盖8类灾害、28个子类,含120+历史极端事件和数千个生命周期问答样本。
- 按预测、追踪、评估三阶段设计评测体系,贴近真实应急响应流程。
多模态大语言模型(MLLM)在解读地球观测数据方面日益重要,但其在真实灾害应急响应中的能力仍缺乏充分评估。现有遥感基准多依赖静态、事后且经专家处理的数据(如网格化再分析数据),难以匹配灾害快速演变、需在严格时限内决策的现实场景。为此,我们提出 Obshazard-bench,一个面向实时灾害智能的观测驱动型基准。该基准直接整合来自多种卫星传感器的原始高频卫星探测流,同步融合地面站观测、历史灾害记录及社会经济指标,绕过延迟的专家处理与物理反演流程。涵盖8大灾害类别、28个子类别,覆盖60多个国家,包含超过120个历史极端事件案例及数千个生命周期导向的视觉问答样本。此外,基准定义了与实际灾害工作流对齐的三阶段评估体系:预测性危机预判(灾前风险识别与早期预警)、动态演化推理(灾中跟踪与终止预测)、多维度影响量化(灾后强度推断、人道负担估算与社会经济影响评估)。对代表性通用与地球聚焦基础模型的实验表明,当前模型在将原始多通道物理观测转化为时间对齐、决策相关的灾害推理方面存在显著局限。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. Existing remote sensing benchmarks largely rely on static, post-hoc, and expert-processed products, such as gridded reanalysis data, which are difficult to align with operational disaster scenarios where hazards evolve rapidly and decisions must be made under strict time constraints. To bridge this gap, we introduce Obshazard-bench, a real-time, observation-driven benchmark for evaluating disaster intelligence in MLLMs. Unlike image-centric or post-event benchmarks, Obshazard-bench directly integrates raw, high-frequency satellite sounding streams from diverse satellite sensors with concurrent ground-station observations, historical disaster records, and socio-economic indicators, bypassing delayed expert-processing and physical-inversion pipelines. The benchmark covers 8 major disaster categories and 28 sub-categories across more than 60 countries, incorporating over 120 historically documented extreme-event cases and thousands of lifecycle-oriented VQA samples. Moreover, Obshazard-bench further defines a three-stage evaluation taxonomy aligned with the operational disaster workflow: Predictive Crisis Anticipation for pre-disaster risk detection and early forecasting, Active Evolution Reasoning for in-situ disaster tracking and termination prediction, and Multi-faceted Impact Quantification for post-disaster magnitude deduction, humanitarian burden estimation, and socio-economic impact assessment. Experiments on representative general-purpose and Earth-focused foundation models reveal substantial limitations in transforming raw multi-channel physical observations into temporally grounded and decision-relevant disaster reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。