构建首个面向火灾监测的热成像多模态问答基准,提升空中火情智能分析能力。
WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring

- 融合可见光与辐射测温热成像,构建6097组多模态样本
- 包含207,298道多选题,覆盖火情检测、定位、分类等六大任务
- 结合大模型生成与人工验证,确保标注可靠性,适合灾害应急研究者使用
火灾监测需依赖机载平台提供及时、可操作的情境感知,但现有航空视觉问答(VQA)基准未针对火灾场景评估基于热成像的多模态推理。本文提出WildFireVQA,一个大规模的航空火灾监测多模态问答基准,整合可见光图像与辐射测温热数据。该基准包含6,097个RGB-热成像样本,每组样本包含一幅RGB图像、一幅伪彩色热图及一幅辐射测温TIFF文件,并配以34个问题,共生成207,298道多选题,涵盖存在性与检测、分类、分布与分割、定位与方向、跨模态推理以及飞行规划等六类任务,服务于实时火灾情报。为提升标注可靠性,采用多模态大语言模型(MLLM)生成答案结合传感器驱动确定性标注、人工校验及帧内/帧间一致性检查。进一步建立了基于辐射测温统计量的综合评估协议,对代表性MLLM在RGB、热成像及检索增强设置下进行评估。实验表明,当前模型中RGB仍为最强模态,而检索增强热数据对强模型带来增益,凸显温度引导推理的价值与现有模型在关键安全场景中的局限性。数据集与代码已开源:https://github.com/mobiiin/WildFire_VQA。
原文摘要 · Abstract (English)
Wildfire monitoring requires timely, actionable situational awareness from airborne platforms, yet existing aerial visual question answering (VQA) benchmarks do not evaluate wildfire-specific multimodal reasoning grounded in thermal measurements. We introduce WildFireVQA, a large-scale VQA benchmark for aerial wildfire monitoring that integrates RGB imagery with radiometric thermal data. WildFireVQA contains 6,097 RGB-thermal samples, where each sample includes an RGB image, a color-mapped thermal visualization, and a radiometric thermal TIFF, and is paired with 34 questions, yielding a total of 207,298 multiple-choice questions spanning presence and detection, classification, distribution and segmentation, localization and direction, cross-modal reasoning, and flight planning for operational wildfire intelligence. To improve annotation reliability, we combine multimodal large language model (MLLM)-based answer generation with sensor-driven deterministic labeling, manual verification, and intra-frame and inter-frame consistency checks. We further establish a comprehensive evaluation protocol for representative MLLMs under RGB, Thermal, and retrieval-augmented settings using radiometric thermal statistics. Experiments show that across task categories, RGB remains the strongest modality for current models, while retrieved thermal context yields gains for stronger MLLMs, highlighting both the value of temperature-grounded reasoning and the limitations of existing MLLMs in safety-critical wildfire scenarios. The dataset and benchmark code are open-source at https://github.com/mobiiin/WildFire_VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。