arXiv:2512.09663cs.CV2025-12被引 7

首个红外图像多模态理解基准,提升模型对热成像的识别能力。

IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting

  • 构建首个红外图像多模态评测基准,覆盖10个理解维度。
  • 提出无需训练的生成式视觉提示方法,显著提升40+模型性能。
  • 适合研究红外视觉、多模态模型或跨域迁移的学者使用。

近年来,多模态大语言模型(MLLMs)在多个基准上取得显著进展,但其对红外图像的理解能力仍未知。为此,我们提出IF-Bench,首个高质量红外图像多模态理解评测基准。该基准包含来自23个红外数据集的499张图像和680个精心设计的图文问答对,覆盖10个关键图像理解维度。基于此,我们系统评估了40余种开源与闭源MLLM,采用循环评估、双语测评和混合判断策略以提高结果可靠性。分析揭示模型规模、架构和推理范式对红外图像理解的影响。此外,我们提出一种无需训练的生成式视觉提示(GenViP)方法,利用先进图像编辑模型将红外图像转换为语义与空间对齐的RGB图像,缓解域分布偏移问题。大量实验表明,该方法在多种MLLM上均带来显著性能提升。基准与代码已公开于https://github.com/casiatao/IF-Bench。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have led to impressive progress across various benchmarks. However, their capability in understanding infrared images remains unexplored. To address this gap, we introduce IF-Bench, the first high-quality benchmark designed for evaluating multimodal understanding of infrared images. IF-Bench consists of 499 images sourced from 23 infrared datasets and 680 carefully curated visual question-answer pairs, covering 10 essential dimensions of image understanding. Based on this benchmark, we systematically evaluate over 40 open-source and closed-source MLLMs, employing cyclic evaluation, bilingual assessment, and hybrid judgment strategies to enhance the reliability of the results. Our analysis reveals how model scale, architecture, and inference paradigms affect infrared image comprehension, providing valuable insights for this area. Furthermore, we propose a training-free generative visual prompting (GenViP) method, which leverages advanced image editing models to translate infrared images into semantically and spatially aligned RGB counterparts, thereby mitigating domain distribution shifts. Extensive experiments demonstrate that our method consistently yields significant performance improvements across a wide range of MLLMs. The benchmark and code are available at https://github.com/casiatao/IF-Bench.

红外图像多模态视觉提示评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。