arXiv:2508.06585cs.AIcs.CV2025-08被引 16

测试大模型在真实场景下的数数能力,发现表现远低于预期。

CountQA: How Well Do MLLMs Count in the Wild?

  • 构建高密度、遮挡复杂的现实图像数据集,评估模型数数能力。
  • 15个主流多模态模型平均准确率仅42.9%,数量越多越差。
  • 适合关注视觉认知与数字理解的研究者,推动模型可靠性提升。

多模态大语言模型虽在理解视觉场景上表现流畅,但在基础认知能力——物体计数上存在明显缺陷,严重限制其在真实应用中的可靠性。现有评测基准或对象密度低,或局限于特定视觉领域,无法充分检验模型在复杂场景下的表现。为此,我们提出CountQA,一个包含超过1,500个问答对的新基准,涵盖高密度、杂乱和遮挡的现实图像。我们评估了15个代表性MLLMs,发现表现最好的模型准确率仅为42.9%,且随着物体数量增加,性能持续下降。CountQA为诊断并修复这一核心缺陷提供了专门评测工具,助力下一代具备数值准确性与空间感知能力的多模态模型发展。数据集与代码将在论文接收后开源。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) demonstrate remarkable fluency in understanding visual scenes, yet they exhibit a critical lack in a fundamental cognitive skill: object counting. This blind spot severely limits their reliability in real-world applications. To date, this capability has been largely unevaluated in complex scenarios, as existing benchmarks either feature sparse object densities or are confined to specific visual domains, failing to test models under realistic conditions. Addressing this gap, we introduce CountQA, a challenging new benchmark designed to probe this deficiency. Comprising over 1,500 question-answer pairs, CountQA features real-world images with high object density, clutter, and occlusion. We investigate this weakness by evaluating 15 prominent MLLMs on the CountQA benchmark and reveal that the top-performing model achieves a mere 42.9% accuracy, with performance declining as object counts rise. By providing a dedicated benchmark to diagnose and rectify this core weakness, CountQA paves the way for a new generation of MLLMs that are not only descriptively fluent but also numerically grounded and spatially aware. We will open-source the dataset and code upon paper acceptance to foster further research.

多模态计数能力评测基准视觉认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。