arXiv:2608.15425cs.CVcs.AI2026-08

构建认知启发的计数基准,揭示视觉语言模型对数量感知的理解差异。

NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

论文配图:NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models
图 1 · 摘自论文原文
  • 设计10,800张合成图像,独立控制大小、排列与数量,逐步消融纹理形状颜色。
  • 模型架构解释32.5%性能差异,早期视觉编码器已出现可分的数量信号。
  • 适合研究多模态模型认知能力、数量理解机制的研究者使用。

视觉语言模型(VLMs)在高层多模态任务中表现优异,但数量感知——一种人类婴儿在语言习得前就具备的认知能力——在当前模型中仍不清晰,因现有计数基准将数量与相关视觉因素混杂。我们提出一个认知启发的诊断性基准NumerosityVLM,包含10,800张合成图像,覆盖六种受控条件。该基准正交地操控对象大小、空间排列与数量,同时逐步消除纹理、形状和颜色。在零样本设置下评估七种VLMs,多因子分析显示模型架构解释了最大比例的性能方差(偏ω²=0.325),远超视觉条件影响。层间探测进一步表明,线性可分的数量信号稳定出现在视觉编码器的早期阶段,而不同模型间的性能差异主要与语言模型组件相关。代码与数据公开于https://github.com/fuy3/NumerosityVLM-Benchmark 和 https://huggingface.co/datasets/fuy3/NumerosityVLM。

原文摘要 · Abstract (English)

Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors. We introduce a cognitively inspired diagnostic benchmark, NumerosityVLM, comprising 10,800 synthetic images across six controlled conditions. The benchmark orthogonally manipulates object size, spatial arrangement, and numerosity, while progressively ablating texture, shape, and color. Evaluating seven VLMs in a zero-shot setting, multi-factor analysis reveals that model architecture explains the largest proportion of performance variance (partial $ω^{2}=0.325$), far exceeding visual conditions. Layer-wise probing further shows that linearly separable numerosity signals consistently emerge at early stages of the vision encoder, while performance differences across evaluated models are primarily associated with the language model component. Code and data are publicly available at https://github.com/fuy3/NumerosityVLM-Benchmark, and https://huggingface.co/datasets/fuy3/NumerosityVLM.

视觉语言模型数量感知认知基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。