构建首个面向灾情地理智能的多智能体评测基准,提升灾害分析准确性。
GeoDisaster: Benchmarking Orchestrated Agents for Operational Disaster Geo-Intelligence

- 设计18个专用工具的多智能体协同框架,通过执行合约实现分工协作。
- 涵盖43类问题、2921个真实灾情实例,覆盖五类地理灾害任务。
- 提出RCEA机制,显著提升工具使用、证据关联与决策一致性。
遥感视觉语言模型在地球观测分析中已实现视觉解读与指令响应,但在需要工具驱动的空间推理和结构化证据决策的灾情地理智能场景中仍显不足。本文提出GeoDisaster,一个包含2,921个经验证实例的灾情地理推理评测基准,覆盖43类问题与五大任务类型:森林砍伐监测、多灾害分析、建筑损毁评估、洪水安全路径规划及哨兵-1 SAR洪水监测。数据融合光学与雷达影像、栅格掩码、矢量几何、路网与暴露层等异构地球观测/地理信息系统(EO/GIS)证据,涵盖灾害检测、损毁评估、暴露估算与诊断报告生成。真值答案基于可执行地理空间工作流与确定性一致性校验,无需依赖语言模型标注。我们进一步提出一种由18个灾情专用工具构成的多智能体协同框架,角色专业化代理通过显式执行合约协调,利用角色-合约期望对齐(RCEA)进行故障感知的监督微调与密集步骤级信号下的合约强化学习。实验表明,GeoDisaster挑战现有遥感视觉语言模型与智能体系统,而RCEA有效提升工具使用、证据锚定、状态一致性和决策生成能力。
原文摘要 · Abstract (English)
Remote-sensing vision-language models (RS-VLMs) have advanced Earth-observation analysis toward visual interpretation and instruction-following, yet fall short of operational geo-intelligence, which demands tool-grounded spatial reasoning and structured, evidence-backed decisions. We introduce GeoDisaster, an operational geospatial disaster reasoning benchmark with 2,921 verified instances across 43 question types and five task families: deforestation monitoring, multi-hazard analysis, building-damage assessment, flood-safe routing, and Sentinel-1 SAR flood monitoring. Instances integrate heterogeneous EO/GIS evidence-optical and SAR imagery, raster masks, vector geometries, road networks, and exposure layers-spanning hazard detection, damage assessment, exposure estimation, and diagnostic report generation. Ground-truth answers are grounded in executable geospatial workflows and deterministic consistency checks, removing the need for language-model annotation. We further propose an orchestrated multi-agent framework with 18 disaster-oriented tools, where role-specialized agents coordinate through explicit execution contracts, aligned via Role-Contract Expectation Alignment (RCEA): failure-aware supervised fine-tuning combined with contract-grounded reinforcement learning over dense step-level signals. Experiments show that GeoDisaster challenges existing RS-VLMs and agentic systems, while RCEA improves tool use, evidence grounding, state consistency, and decision generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。