arXiv:2606.27128cs.CVcs.RO2026-06

构建首个基于热辐射的无人机火灾视觉问答基准,提升灾情推理可靠性。

FlameVQA: A Physically-Grounded UAV Wildfire VQA Benchmark with Radiometric Thermal Supervision

论文配图:FlameVQA: A Physically-Grounded UAV Wildfire VQA Benchmark with Radiometric Thermal Supervision
图 1 · 摘自论文原文
  • 融合可见光与热红外图像,实现温度感知的多模态火灾分析。
  • 每图34题覆盖6类任务,包含烟雾下火情检测与覆盖估算等挑战性问题。
  • 适用于灾害监测、无人机智能系统研发人员参考使用。

从无人机视角进行野火监测需要对复杂空域场景进行可靠推理,而烟雾、尺度变化和遮挡常限制仅依赖可见光图像的解读。本文提出FlameVQA,一个基于FLAME 3数据集的无人机火灾视觉问答基准,利用配对的RGB图像与辐射测温热红外TIFF图像,实现基于温度信息的安全关键型推理。FlameVQA包含每张图像34道多选题,涵盖六类操作能力:检测、定位、分布/覆盖估算、跨模态推理及飞行规划。为确保标签可靠性,采用多模态大模型辅助标注结合确定性热力学规则与跨问题一致性检查,并经人工审核。同时评估代表性多模态大模型在该基准上的表现,提供未来研究基线。结果表明,当存在明确跨模态线索时模型表现良好,但在重烟环境下火源检测和覆盖估算上仍存在显著失败。这表明当前多模态大模型需进行领域特定适应以更好支持灾情与野火监测。数据集与评测代码已开源至github.com/mobiiin/WildFire_VQA。

原文摘要 · Abstract (English)

Wildfire monitoring from UAVs requires reliable reasoning over complex aerial scenes, where smoke, scale variation, and occlusions often limit RGB-only interpretation. We introduce FlameVQA, a multiple-choice visual question answering benchmark for UAV-based wildfire intelligence built on FLAME 3, leveraging paired RGB imagery and radiometric thermal TIFFs for temperature-grounded, safety-critical reasoning. FlameVQA includes 34 multiple-choice questions per image spanning six operational capability groups, covering tasks such as detection, localization, distribution/coverage estimation, cross-modal reasoning, and flight planning. To ensure label reliability, we combine MLLM-assisted annotation with deterministic thermal rules and cross-question consistency checks, followed by human auditing. We also evaluate representative MLLMs on FlameVQA to provide baselines for future work. Results show strong performance when explicit cross-modal cues are available, but notable failures on presence detection under heavy smoke and on coverage estimation. These findings suggest that current MLLMs require domain-specific adaptation to better support disaster and wildfire monitoring. The dataset and benchmark code are open-source at github.com/mobiiin/WildFire_VQA

火灾监测视觉问答无人机多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。