arXiv:2512.11215cs.CV2025-12中稿 · WACV 2026被引 4

评测大模型识别野火烟雾能力,发现早期定位仍困难

SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection

  • 构建四个任务的SmokeBench基准,评估多模态大模型烟雾识别与定位能力
  • 大模型对大面积烟雾分类有效,但早期烟雾定位准确率普遍偏低
  • 烟雾体积影响模型表现,对比度影响较小,适合火灾监测研究者参考

野火烟雾透明、无定形,常与云层视觉混淆,早期检测尤为困难。本文提出SmokeBench基准,用于评估多模态大语言模型(MLLMs)在图像中识别和定位野火烟雾的能力。该基准包含四项任务:(1)烟雾分类,(2)基于瓦片的烟雾定位,(3)基于网格的烟雾定位,(4)烟雾检测。我们评估了Idefics2、Qwen2.5-VL、InternVL3、Unified-IO 2、Grounding DINO、GPT-4o和Gemini-2.5 Pro等多款MLLMs。结果表明,部分模型在烟雾覆盖面积较大时可准确分类,但所有模型在早期阶段的定位均表现不佳。进一步分析显示,烟雾体积与模型性能强相关,而对比度影响较小。这些发现揭示了当前MLLMs在安全关键型野火监测中的显著局限性,凸显了提升早期烟雾定位方法的迫切需求。

原文摘要 · Abstract (English)

Wildfire smoke is transparent, amorphous, and often visually confounded with clouds, making early-stage detection particularly challenging. In this work, we introduce a benchmark, called SmokeBench, to evaluate the ability of multimodal large language models (MLLMs) to recognize and localize wildfire smoke in images. The benchmark consists of four tasks: (1) smoke classification, (2) tile-based smoke localization, (3) grid-based smoke localization, and (4) smoke detection. We evaluate several MLLMs, including Idefics2, Qwen2.5-VL, InternVL3, Unified-IO 2, Grounding DINO, GPT-4o, and Gemini-2.5 Pro. Our results show that while some models can classify the presence of smoke when it covers a large area, all models struggle with accurate localization, especially in the early stages. Further analysis reveals that smoke volume is strongly correlated with model performance, whereas contrast plays a comparatively minor role. These findings highlight critical limitations of current MLLMs for safety-critical wildfire monitoring and underscore the need for methods that improve early-stage smoke localization.

多模态模型野火监测烟雾检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。