arXiv:2603.07888cs.CVcs.AI2026-03被引 4

构建细粒度视觉对比推理基准,揭示VLM与人类差距

VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?

  • 设计涵盖10类细微差异的跨领域图像对比任务
  • 实测多类VLM在医疗/工业等场景中表现显著落后于人类
  • 首次系统分析模型在细粒度推理中的失效边界

区分视觉相似图像间的细微差异对工业缺陷检测、医学影像和航空监视等应用至关重要。尽管现有视觉语言模型(VLM)比较推理基准已出现,但多聚焦于显著差异图像,难以反映真实场景所需的精细推理能力。本文提出VLM-SubtleBench,一个评估VLM在细微比较推理方面表现的基准。该基准涵盖10类差异类型:属性、状态、情绪、时间、空间、存在性、数量、质量、视角和动作,并构建了反映这些细粒度变化的配对问题-图像数据集。不同于以往仅限自然图像数据集的基准,本基准覆盖工业、航空和医学等多种领域。通过对多种专有及开源VLM的广泛评估,我们发现模型在各类差异类型和不同领域中均存在系统性性能差距,且通过受控分析揭示了模型推理能力急剧下降的关键场景。该基准与研究结果共同为推动VLM向人类级比较推理迈进奠定了基础。

原文摘要 · Abstract (English)

The ability to distinguish subtle differences between visually similar images is essential for diverse domains such as industrial anomaly detection, medical imaging, and aerial surveillance. While comparative reasoning benchmarks for vision-language models (VLMs) have recently emerged, they primarily focus on images with large, salient differences and fail to capture the nuanced reasoning required for real-world applications. In this work, we introduce VLM-SubtleBench, a benchmark designed to evaluate VLMs on subtle comparative reasoning. Our benchmark covers ten difference types - Attribute, State, Emotion, Temporal, Spatial, Existence, Quantity, Quality, Viewpoint, and Action - and curate paired question-image sets reflecting these fine-grained variations. Unlike prior benchmarks restricted to natural image datasets, our benchmark spans diverse domains, including industrial, aerial, and medical imagery. Through extensive evaluation of both proprietary and open-source VLMs, we reveal systematic gaps between model and human performance across difference types and domains, and provide controlled analyses highlighting where VLMs' reasoning sharply deteriorates. Together, our benchmark and findings establish a foundation for advancing VLMs toward human-level comparative reasoning.

视觉语言模型细粒度推理基准测试医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。