构建多模态灾难评估基准,推动遥感图像理解向真实救援需求迈进
DisasterInsight: A Multimodal Benchmark for Function-Aware and Grounded Disaster Assessment
- 基于xbd数据重构11.2万栋建筑级实例,支持多任务指令评测
- 模型在损伤程度与报告生成上表现不佳,功能分类仍具挑战
- 适配人道主义评估流程,适合灾难响应与遥感智能研究者使用
及时解读卫星图像对灾后响应至关重要,但现有遥感视觉-语言基准多关注粗粒度标签和图像级识别,忽视真实人道主义工作流所需的语义功能理解与指令鲁棒性。本文提出DisasterInsight,一个面向真实灾难分析任务的多模态评估基准。该基准将xBD数据集重构为约11.2万栋建筑为中心的实例,支持多任务、多指令的评估,涵盖建筑功能分类、损毁等级与灾种分类、计数及符合人道评估指南的结构化报告生成。为建立领域适配基线,提出DI-Chat:通过低秩适应(LoRA)在灾害特定指令数据上微调现有视觉-语言模型骨干网络。对先进通用与遥感专用模型的大量实验表明,各模型在损伤理解与报告生成任务中存在显著差距,而DI-Chat在损毁等级与灾种分类、报告质量上取得显著提升,建筑功能分类仍为所有模型的难点。DisasterInsight为研究灾难影像中的具身多模态推理提供统一基准。
原文摘要 · Abstract (English)
Timely interpretation of satellite imagery is critical for disaster response, yet existing vision-language benchmarks for remote sensing largely focus on coarse labels and image-level recognition, overlooking the functional understanding and instruction robustness required in real humanitarian workflows. We introduce DisasterInsight, a multimodal benchmark designed to evaluate vision-language models (VLMs) on realistic disaster analysis tasks. DisasterInsight restructures the xBD dataset into approximately 112K building-centered instances and supports instruction-diverse evaluation across multiple tasks, including building-function classification, damage-level and disaster-type classification, counting, and structured report generation aligned with humanitarian assessment guidelines. To establish domain-adapted baselines, we propose DI-Chat, obtained by fine-tuning existing VLM backbones on disaster-specific instruction data using parameter-efficient Low-Rank Adaptation (LoRA). Extensive experiments on state-of-the-art generic and remote-sensing VLMs reveal substantial performance gaps across tasks, particularly in damage understanding and structured report generation. DI-Chat achieves significant improvements on damage-level and disaster-type classification as well as report generation quality, while building-function classification remains challenging for all evaluated models. DisasterInsight provides a unified benchmark for studying grounded multimodal reasoning in disaster imagery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。