构建多模态大模型计数能力的综合评测基准,揭示其在复杂场景下的定量缺陷。
HoloCount: A Holistic Visual Counting Benchmark for MLLMs

- 设计三级分层评测体系,涵盖语义计数、逻辑分析与鲁棒性测试。
- 超过20个顶尖模型在复杂推理任务中表现显著下降,最高误差超50%。
- 适合研究多模态模型可靠性、计数精度及对抗性测试的学者使用。
视觉计数是多模态智能的基础,需精细定位与空间推理的无缝融合。尽管多模态大语言模型(MLLMs)在定性场景理解上取得显著进展,其定量精度仍存在明显瓶颈,常出现持续性的数值幻觉。现有计数评测主要聚焦简化情境下的基础感知,无法捕捉逻辑约束或对抗条件下的复杂失效模式。为此,我们提出HoloCount,一个结构化的全景式评测基准,采用三级层次化分类:(1) 语义计数,关注原子与属性级枚举;(2) 分析计数,评估基于空间与集合推理的逻辑组合;(3) 鲁棒性测试,探测高密度场景、语言偏见等对抗情形下的模型完整性。对超过20个先进MLLMs的全面评估显示,即使顶级模型在从感知过渡到复杂分析与恶劣场景时性能显著退化。研究结果系统描绘了当前MLLM计数能力图景,并为构建更可靠、具身的多模态系统提供了路线图。数据集可访问:https://mm-mvr.github.io/HoloCount/。
原文摘要 · Abstract (English)
Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Language Models (MLLMs) have achieved remarkable success in qualitative scene understanding, their quantitative precision remains a significant bottleneck, often characterized by persistent numerical hallucinations. Existing counting benchmarks primarily focus on basic perception in simplified contexts, failing to capture the complex failure modes that emerge under logical constraints or adversarial conditions. To address these limitations, we introduce HoloCount, a holistic and diagnostically rich benchmark structured around a three-level hierarchical taxonomy. HoloCount evaluates MLLMs across: (1) Semantic Counting, focusing on atomic and property-based enumeration; (2) Analytical Counting, assessing logical composition through spatial and set-based reasoning; and (3) Robustness Testing, probing model integrity against adverse scenarios and grounded counter-priors, such as high-density scenes and linguistic biases. Through an exhaustive evaluation of over 20 state-of-the-art MLLMs, we reveal a critical performance gap: even top-tier models degrade significantly as tasks transition from perception to complex analytical reasoning and adverse scenarios. Our findings provide a systematic landscape of current MLLM counting capabilities and offer a roadmap for developing more grounded and reliable multimodal systems. The dataset is available at https://mm-mvr.github.io/HoloCount/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。