arXiv:2602.14989cs.CVcs.AI2026-02KDD被引 2

首个针对热成像视觉语言模型的系统性评测基准

ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery

  • 构建包含5.5万条热成像图文问答的数据集,覆盖室内外多场景
  • 发现主流模型在温度推理上表现差,对颜色映射变换敏感
  • 适合研究热成像理解、多模态感知与鲁棒性评估的学者

视觉语言模型(VLMs)在可见光图像上表现优异,但在热成像上泛化能力不足。热成像在夜间监控、搜救、自动驾驶和医疗筛查等光照缺失场景中至关重要。与可见光图像不同,热成像反映的是物理温度而非颜色或纹理,需具备感知与推理能力,而现有以可见光为中心的评测基准无法有效评估。我们提出ThermEval-B,一个约5.5万条热成像图文问答的结构化基准,用于评估热视觉语言理解所需的基础能力。该基准整合公开数据集与我们新收集的ThermEval-D,这是首个提供密集像素级温度图及跨室内外环境的语义部位标注的数据集。评估25个开源与闭源VLMs发现,模型在温度相关推理上持续失败,对颜色映射变换敏感,倾向于依赖语言先验或固定回应,提示词或微调仅带来微弱提升。结果表明,热成像理解需超越可见光中心假设的专用评测,ThermEval为推动热视觉语言建模发展提供关键基准。

原文摘要 · Abstract (English)

Vision language models (VLMs) achieve strong performance on RGB imagery, but they do not generalize to thermal images. Thermal sensing plays a critical role in settings where visible light fails, including nighttime surveillance, search and rescue, autonomous driving, and medical screening. Unlike RGB imagery, thermal images encode physical temperature rather than color or texture, requiring perceptual and reasoning capabilities that existing RGB-centric benchmarks do not evaluate. We introduce ThermEval-B, a structured benchmark of approximately 55,000 thermal visual question answering pairs designed to assess the foundational primitives required for thermal vision language understanding. ThermEval-B integrates public datasets with our newly collected ThermEval-D, the first dataset to provide dense per-pixel temperature maps with semantic body-part annotations across diverse indoor and outdoor environments. Evaluating 25 open-source and closed-source VLMs, we find that models consistently fail at temperature-grounded reasoning, degrade under colormap transformations, and default to language priors or fixed responses, with only marginal gains from prompting or supervised fine-tuning. These results demonstrate that thermal understanding requires dedicated evaluation beyond RGB-centric assumptions, positioning ThermEval as a benchmark to drive progress in thermal vision language modeling.

热成像多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。