构建多模态大模型火灾烟雾理解的高精度评测基准
SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

- 设计涵盖20种真实场景的83K图像与193K多选题评测集
- 开源模型平均准确率仅61.9%,暴露安全推理短板
- 用7%领域数据微调,分类准确率提升至64.5%
多模态大语言模型在视觉-语言任务上进展迅速,但在安全关键场景下的可靠性仍待检验。火灾烟雾理解对公共安全与灾害应对至关重要,但现有评测基准缺乏多样化的现实场景和上下文感知评估。我们提出SAFIRE,一个大规模火灾烟雾理解评测基准,包含来自20个场景的83,000张带标注图像,以及从9,700张图像子集中生成的193,000道多选题视觉问答(MCVQA),覆盖从基础感知到高级推理共10个评估维度。通过GPT-5.4辅助的多阶段验证流程与多模态大模型多数投票确保标注质量。对10个开源模型(8B-38B)的评估显示平均准确率为61.9%,暴露出安全关键推理的重大差距。进一步实验表明,仅用7%领域数据微调视觉编码器,即可将火灾场景分类准确率从20.1%提升至64.5%,说明精心构建的数据即使量少也能带来显著性能提升。所有数据集、模型与代码均已公开。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware evaluation. We introduce SAFIRE, a large-scale benchmark for fire-smoke understanding in MLLMs, comprising 83K captioned images from 20 scenarios and 193K multiple-choice VQA (MCVQA) generated from a 9.7K-image subset, spanning 10 evaluation dimensions from basic perception to higher-order reasoning. A GPT-5.4-assisted multi-stage verification pipeline with MLLM majority voting ensures annotation quality. Evaluating ten open-source MLLMs (8B-38B) yields an average accuracy of 61.9%, exposing major gaps in safety-critical reasoning. We further show that adapting vision encoders with only 7% of our domain-specific data boosts fire-scene classification accuracy from 20.1% to 64.5%, indicating that carefully curated data can yield substantial gains even when data volume is limited. All datasets, models, and code are available at https://risys-lab.github.io/SAFIRE/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。