首个评估视觉语言模型理解红温图像对的基准,揭示模型在热成像任务上的严重短板。
RGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models
- 构建14项技能维度的红温图像问答基准,含1600+专家标注问题。
- 采用双精度指标,发现最强模型在热成像理解上仍表现不佳。
- 适合关注多模态鲁棒性、红外视觉理解的研究者使用。
我们提出RGB-Th-Bench,首个专门评估视觉语言模型(VLMs)理解红温图像对能力的基准。尽管VLMs在可见光视觉推理中取得显著进展,但其评估仍主要局限于基于RGB的基准,导致红外视觉任务评估存在关键空白。现有可见-红外数据集或任务专用,或缺乏高质量标注以支撑严格模型评估。RGB-Th-Bench提供了覆盖14个不同技能维度的全面评估框架,包含1,600+专家标注的二选一问答。该基准采用两种准确率指标:标准问题级准确率和更严格的技能级准确率,后者评估模型在每项技能维度下多个问题中的鲁棒性。此设计确保对模型性能的全面评估,包括对抗性和幻觉响应的抵御能力。我们在19个前沿VLM上进行了广泛评估,揭示了在红温图像理解方面存在显著性能差距。结果显示,即使最强模型在热成像理解上仍表现受限,其性能严重依赖于其基于RGB的能力。此外,预训练阶段缺乏大规模、特定应用且经专家标注的热成像-描述对数据集,是观察到性能差距的重要原因。RGB-Th-Bench凸显了推进多模态学习以弥合可见光与热成像理解差距的紧迫性。数据集可通过此链接获取,评估代码也将公开。
原文摘要 · Abstract (English)
We introduce RGB-Th-Bench, the first benchmark designed to evaluate the ability of Vision-Language Models (VLMs) to comprehend RGB-Thermal image pairs. While VLMs have demonstrated remarkable progress in visual reasoning and multimodal understanding, their evaluation has been predominantly limited to RGB-based benchmarks, leaving a critical gap in assessing their capabilities in infrared vision tasks. Existing visible-infrared datasets are either task-specific or lack high-quality annotations necessary for rigorous model evaluation. To address these limitations, RGB-Th-Bench provides a comprehensive evaluation framework covering 14 distinct skill dimensions, with a total of 1,600+ expert-annotated Yes/No questions. The benchmark employs two accuracy metrics: a standard question-level accuracy and a stricter skill-level accuracy, which evaluates model robustness across multiple questions within each skill dimension. This design ensures a thorough assessment of model performance, including resilience to adversarial and hallucinated responses. We conduct extensive evaluations on 19 state-of-the-art VLMs, revealing significant performance gaps in RGB-Thermal understanding. Our results show that even the strongest models struggle with thermal image comprehension, with performance heavily constrained by their RGB-based capabilities. Additionally, the lack of large-scale application-specific and expert-annotated thermal-caption-pair datasets in pre-training is an important reason of the observed performance gap. RGB-Th-Bench highlights the urgent need for further advancements in multimodal learning to bridge the gap between visible and thermal image understanding. The dataset is available through this link, and the evaluation code will also be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。