arXiv:2605.11633cs.AI2026-05被引 4

首个端到端灾情响应智能体评测基准,检验大模型在真实灾害场景下的综合决策能力。

Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations

论文配图:Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations
图 1 · 摘自论文原文
  • 构建涵盖45起真实灾害的515个任务,模拟从感知到报告全流程
  • 发现大模型在灾情语义对齐和工具调用顺序上存在系统性缺陷
  • 适合研究应急智能体、地理空间推理与多模态决策的学者使用

灾情应急响应不仅涉及损毁评估,还需整合多源传感信号,推理道路网络、人口分布与关键设施,规划疏散并生成可行动报告。现有工作多局限于遥感感知或通用工具使用,缺乏对应急全流程的端到端评估。本文提出灾情应急响应智能体基准(DORA),首个面向全流程灾情响应的智能体评测基准:包含45起真实灾害事件的515个专家撰写任务,配套总计3,500步工具调用的专家验证可复现黄金轨迹。任务覆盖灾情感知、空间关系分析、救援疏散规划、时序演化推理及多模态报告合成五个维度。智能体通过108个工具组成的MCP库,在异构地理空间数据(光学、SAR、多光谱影像,0.015-10m GSD)与高程、社会矢量层上操作。我们全面评估13个前沿大模型,揭示三大持续挑战:1)灾情领域对齐存在独特失效模式(损毁语义对齐、传感器模态错配、流程组合错误);2)工具选择与参数定位双重瓶颈,即使提供黄金调用顺序提示,准确率提升仅1.08%-4.40%,其他辅助结构最多提升3.24%;3)组合脆弱性随流程长度加剧,长流程中智能体与黄金轨迹差距从7%扩大至56%。DORA为构建可操作可靠的灾情响应智能体提供了严格测试平台。

原文摘要 · Abstract (English)

Operational disaster response goes beyond damage assessment, requiring responders to integrate multi-sensor signals, reason over road networks, populations and key facilities, plan evacuations, and produce actionable reports. However, prior work largely isolates remote-sensing perception or evaluates generic tool use, leaving the end-to-end workflows of emergency operations underexplored. In this paper, we introduce Disaster Operational Response Agent benchmark (DORA), the first agentic benchmark for end-to-end disaster response: 515 expert-authored tasks across 45 real-world disaster events spanning 10 types, paired with expert-verified, replayable gold trajectories totaling 3,500 tool-call steps. Tasks span five dimensions that cover the operational disaster-response pipeline: disaster perception, spatial relational analysis, rescue and evacuation planning, temporal evolution reasoning, and multi-modal report synthesis. Agents compose calls from a 108-tool MCP library over heterogeneous geospatial data: optical, SAR, and multi-spectral imagery across single-, bi-, and multi-temporal sequences (0.015-10m GSD), complemented by elevation and social vector layers. We comprehensively evaluate 13 frontier LLMs on our benchmark, revealing three persistent challenges: 1) disaster-domain grounding exposes unique failure modes (damage-semantic grounding, sensor-modality mismatch, and disaster-pipeline composition); 2) agents are doubly bottlenecked by tool selection and argument grounding, where gold tool-order hints improve accuracy by only 1.08-4.40%, and alternative scaffolds yield at most a 3.24% gain; 3) compositional fragility scales with trajectory length, the agent-to-gold gap widening from 7% to 56% on long pipelines. DORA establishes a rigorous testbed for operationally reliable disaster-response agents.

灾情响应智能体评测地理空间推理大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。