揭示视觉语言模型在数量识别上的关键瓶颈,发现其符号映射环节失效。
Unveiling the Visual Counting Bottleneck in Vision-Language Models
- 拆解视觉计数为三阶段:个体识别、数量感知、符号映射。
- 模型能感知数量但无法正确命名,尤其在未见过的数量上失败。
- 适合研究模型泛化能力与跨模态对齐的学者参考。
尽管大型视觉语言模型(VLMs)在插值任务中表现优异,但在系统性外推时出现灾难性失败,尤其体现在视觉计数上。本文将视觉计数分解为三个认知阶段:视觉个体化、数量意识与符号映射。通过合成围棋盘和线性探针分析,我们发现视觉主干网络在超出训练范围的计数任务中仍保持稳定的线性可分表示,排除了感知层面的失败。此外,模型保留了潜在的数量意识,能对未计数的数量进行比较推理。问题出在符号映射阶段,模型无法将视觉数量有效映射到符号标记。研究支持‘断裂数量假设’:VLMs 未学习统一的数域空间,而是习得分离的、模态特定的统计流形,导致跨模态对齐失败。该结论在当前最先进的基础模型上得到验证,表明仅靠数据扩展不足以弥补此缺陷,需引入归纳偏置以建立统一表征。
原文摘要 · Abstract (English)
While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting. In this work, we investigate this extrapolation bottleneck by deconstructing visual counting into three cognitive stages: visual individuation, magnitude awareness, and symbolic mapping. Using synthetic Go boards and linear probes, we demonstrate that visual backbones maintain robust, linearly separable representations of quantity well into the extrapolation regime, ruling out perceptual failure. Furthermore, models retain latent magnitude awareness, successfully performing comparative reasoning on quantities they fail to enumerate. We pinpoint the collapse to the symbolic mapping stage, where the model fails to project valid visual magnitudes onto symbolic tokens. Our findings support a frac tured magnitude hypothesis: VLMs fail to acquire a universal number space, instead learning disjoint, modality-specific statistical manifolds that prevent cross-modal grounding for unseen quantities. Validated on the state-of-the-art foundation model, our results suggest that bridging this gap requires inductive priors enforcing unified representations, as data scaling alone is insufficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。