测试大模型对细微视觉细节的感知能力,发现表现远低于人类。
HueManity: Probing Fine-Grained Visual Perception in MLLMs
- 用8万多张色盲检测图测试模型图案识别能力。
- 最强模型在数字任务上仅33.6%准确率,远低于人类99.38%。
- 适合关注视觉感知可靠性的研究人员和应用开发者。
近期多模态大语言模型(MLLMs)在视觉问答、图像描述等高层视觉推理任务中表现出色,但现有评测基准普遍忽视其对细微感知细节的捕捉能力。随着MLLMs在安全与可靠性关键场景中的广泛应用,视觉敏锐度变得至关重要。本文提出HueManity,一个可扩展的自动化评测基准,用于评估MLLMs在细粒度视觉感知方面的能力。HueManity包含83,850张类伊希加拉风格图像,嵌入数字和字母字符串,用于评测模式识别这一核心视觉理解能力。我们对九个顶尖MLLMs的评估揭示了显著性能差距:最强模型在简单数字任务上仅达33.6%准确率,在更难的字母数字任务上仅为3%,而人类表现接近满分(99.38%、93.25%),微调后的ResNet-50则分别达到96.5%和94.5%。这些结果暴露了MLLMs在感知基础方面的关键缺陷,该问题被传统侧重高层语义的评测所掩盖。
原文摘要 · Abstract (English)
Recent Multimodal Large Language Models (MLLMs) demonstrate strong high-level visual reasoning on tasks such as visual question answering and image captioning. Yet existing benchmarks largely overlook their ability to capture fine-grained perceptual details. As MLLMs are increasingly deployed in safety and reliability critical settings, perceptual acuity becomes essential. We present HueManity, a scalable automated benchmark for assessing fine-grained visual perception in MLLMs. HueManity comprises 83,850 Ishihara-style images embedding alphanumeric strings, designed to evaluate pattern recognition, a core aspect of visual understanding. Our evaluation of nine state-of-the-art MLLMs uncovers a striking performance deficit: the strongest model achieved only 33.6% accuracy on a simple numeric task and 3% on a harder alphanumeric task, compared to near-ceiling performance from humans (99.38%, 93.25%) and a fine-tuned ResNet-50 (96.5%, 94.5%). These findings expose a critical weakness in MLLMs' perceptual grounding, one that remains obscured by conventional benchmarks emphasizing high-level semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。