arXiv:2606.06217cs.CVcs.AI2026-06

构建无人机灾害响应多阶段推理基准,支持灾前到灾后全流程决策。

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments

论文配图:DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments
图 1 · 摘自论文原文
  • 设计多阶段任务框架,覆盖14类灾情与9项关键响应任务。
  • 提出轻量级DisasterVL模型,在20个任务上接近GPT-4o性能且效率更高。
  • 专为边缘计算优化,适合应急场景下实时推理需求。

灾难发生时,救援人员需判断事件现状、原因、未来趋势及应对策略,常依赖噪声大、视角低的无人机图像,并受现场算力限制。现有跨模态基准多聚焦感知任务,灾种覆盖有限,难以支撑实际应急中的多阶段推理。本文提出DisasterBench,一个面向复杂环境下无人机灾害响应的多阶段跨模态推理基准,涵盖14类灾情场景与9项灾前、灾中、灾后关键任务,包含细粒度灾情-任务映射,明确测试因果归因、扩散预测、损毁分析与决策导向推理能力。为支持边缘部署,提出轻量级多模态模型DisasterVL,采用三阶段训练流程:领域指令微调、思维链引导对齐、强化学习策略优化。在21个主流多模态大模型上的实验表明,20亿参数的DisasterVL优于所有开源模型,显著缩小与闭源顶尖模型差距,实现与GPT-4o相当的推理准确率且更具效率。

原文摘要 · Abstract (English)

When a disaster unfolds, responders must answer not only what is happening, but also why it is happening, what will happen next, and what to do now, often from noisy low-altitude UAV views and under tight on-site compute constraints. However, most existing multimodal benchmarks emphasize perception (e.g., recognition/description), cover limited disaster types, and provide insufficient support for the multi-stage reasoning required in practical emergency response. We introduce DisasterBench, a multi-stage multimodal reasoning benchmark for UAV-Based disaster response in complex environments. DisasterBench spans 14 disaster-related scene types and 9 response-critical tasks across pre-, during-, and post-disaster stages, with fine-grained disaster-task mappings that explicitly test causal attribution, propagation prediction, damage analysis, and decision-oriented reasoning. To enable reasoning on the edge, we further propose DisasterVL, a lightweight multimodal model optimized with a three-stage pipeline combining domain instruction tuning, chain-of-thought-guided multimodal alignment, and reinforcement learning-based policy optimization. Experiments across 21 popular MLLMs show that our 2B-parameter DisasterVL outperforms all evaluated open-source models and substantially narrows the gap to state-of-the-art closed-source models, achieving GPT-4o-comparable reasoning accuracy with superior efficiency. The project page is available at https://github.com/TanmouTT/DisasterBench.

灾害响应多模态推理无人机边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。