测试发现视觉语言模型在组合计数时严重失准,暴露其根本缺陷。
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
- 用简单几何图形构建可控计数任务,隔离干扰因素。
- 单一形状可准确计数,但多种形状组合时错误率显著上升。
- 适合关注VLM推理能力局限的研究者与开发者。
视觉语言模型(VLMs)因其在大规模网络图像-文本数据上训练而备受关注,展现出强大的图像理解、视频理解、复杂视觉推理和具身智能表现。然而,一个基本问题仍待解答:它们能否正确计数?本文提出一个简洁有效的基准测试 VLMCountBench,采用极简设置,仅使用三角形、圆形等基本几何形状及其组合,专注于计数任务,排除其他干扰因素。通过严格控制独立变量,系统研究颜色、大小和提示优化对计数的影响。实验结果表明,当仅存在一种形状类型时,VLMs 能可靠计数;但在多种形状组合时(即组合计数),表现出显著失败,揭示了当前 VLMs 的根本性实证局限,并为未来研究指明重要方向。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have become a central focus of today's AI community, owing to their impressive abilities gained from training on large-scale vision-language data from the Web. These models have demonstrated strong performance across diverse tasks, including image understanding, video understanding, complex visual reasoning, and embodied AI. Despite these noteworthy successes, a fundamental question remains: Can VLMs count objects correctly? In this paper, we introduce a simple yet effective benchmark, VLMCountBench, designed under a minimalist setting with only basic geometric shapes (e.g., triangles, circles) and their compositions, focusing exclusively on counting tasks without interference from other factors. We adopt strict independent variable control and systematically study the effects of simple properties such as color, size, and prompt refinement in a controlled ablation. Our empirical results reveal that while VLMs can count reliably when only one shape type is present, they exhibit substantial failures when multiple shape types are combined (i.e., compositional counting). This highlights a fundamental empirical limitation of current VLMs and motivates important directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。