构建灾难场景视觉问答数据集,评估AI在灾情判断中的实际能力
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes

- 基于真实灾情图像与专家标注问题,覆盖洪水、火灾等多类灾害
- 模型在二元问题上表现好,但对计数和上下文理解仍严重不足
- 适合研究灾难AI、人道主义救援系统及视觉语言模型可靠性
社交媒体图像在自然灾害和人为灾害期间提供了低延迟的情境信息,有助于快速评估损失并响应。尽管视觉问答(VQA)在通用领域表现出色,但在灾情应对所需的复杂且高安全性的推理任务中是否适用尚不明确。我们提出DisasterVQA,一个面向危机情境感知与推理的基准数据集。该数据集包含1,395张真实世界图像和4,405对专家标注的问题-答案对,涵盖洪水、野火、地震等多种灾害事件。数据基于联邦紧急事务管理局(FEMA ESF)和联合国人道主义事务协调办公室(OCHA MIRA)框架设计,包含二元、多选和开放式问题,覆盖态势感知与操作决策任务。我们测试了七种先进视觉语言模型,发现其性能在问题类型、灾害类别、地区和人道任务间存在显著差异。虽然模型在二元问题上表现良好,但在细粒度定量推理、物体计数和上下文敏感解释方面表现不佳,尤其在少数灾害场景下更为明显。DisasterVQA为开发更稳健、具有实际意义的灾情视觉语言模型提供了挑战性且实用的基准。数据集已公开,可访问https://doi.org/10.5281/zenodo.18267769。
原文摘要 · Abstract (English)
Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage assessment and response. While Visual Question Answering (VQA) has shown strong performance in general-purpose domains, its suitability for the complex and safety-critical reasoning required in disaster response remains unclear. We introduce DisasterVQA, a benchmark dataset designed for perception and reasoning in crisis contexts. DisasterVQA consists of 1,395 real-world images and 4,405 expert-curated question-answer pairs spanning diverse events such as floods, wildfires, and earthquakes. Grounded in humanitarian frameworks including FEMA ESF and OCHA MIRA, the dataset includes binary, multiple-choice, and open-ended questions covering situational awareness and operational decision-making tasks. We benchmark seven state-of-the-art vision-language models and find performance variability across question types, disaster categories, regions, and humanitarian tasks. Although models achieve high accuracy on binary questions, they struggle with fine-grained quantitative reasoning, object counting, and context-sensitive interpretation, particularly for underrepresented disaster scenarios. DisasterVQA provides a challenging and practical benchmark to guide the development of more robust and operationally meaningful vision-language models for disaster response. The dataset is publicly available at https://doi.org/10.5281/zenodo.18267769.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。