对比视觉语言模型与专用计数模型的计数能力,发现前者在多数场景下表现更优。
Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
- 用视觉语言模型生成物体位置和名称作为中间提示,提升计数准确率。
- 在两个主流数据集和新构建的精细控制数据集上,多数VLM性能优于专用架构。
- 复杂场景仍无法可靠计数,适合关注开放集计数与提示工程的研究者。
视觉场景中的物品计数是计算机视觉中的基础但具有挑战性的任务。传统方法依赖于特定领域的计数架构,使用预定义类别标注的数据集进行训练。然而,近年来大规模多模态视觉语言模型(VLMs)的发展表明,这些通用架构可能为开放集物品计数提供灵活替代方案。本研究系统比较了最先进的专用计数架构与VLMs在两个流行计数数据集及一个新基准上的表现,该基准对测试图像的视觉属性有更细粒度的控制。结果表明,大多数VLM能够近似估算视觉场景中的物品数量,其表现可媲美甚至超过专用计算机视觉架构。值得注意的是,当VLM被提示生成每个待计数物体的中间表示(如位置和语义标签)时,计数准确率显著提升。然而,所有模型在复杂视觉场景中仍无法可靠计数,说明仍需进一步研究以实现真实环境中可靠的计数系统部署。
原文摘要 · Abstract (English)
Counting the number of items in a visual scene remains a fundamental yet challenging task in computer vision. Traditional approaches to solving this problem rely on domain-specific counting architectures, which are trained using datasets annotated with a predefined set of object categories. However, recent progress in creating large-scale multimodal vision-language models (VLMs) suggests that these domain-general architectures may offer a flexible alternative for open-set object counting. In this study, we therefore systematically compare the performance of state-of-the-art specialized counting architectures against VLMs on two popular counting datasets, as well as on a novel benchmark specifically created to have a finer-grained control over the visual properties of test images. Our findings show that most VLMs can approximately enumerate the number of items in a visual scene, matching or even surpassing the performance of specialized computer vision architectures. Notably, enumeration accuracy significantly improves when VLMs are prompted to generate intermediate representations (i.e., locations and verbal labels) of each object to be counted. Nevertheless, none of the models can reliably count the number of objects in complex visual scenes, showing that further research is still needed to create AI systems that can reliably deploy counting procedures in realistic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。