现有计数模型常误读文本提示,新框架与数据集可更真实评估其语义理解能力。
Does it Really Count? Assessing Semantic Grounding in Text-Guided Class-Agnostic Counting

- 设计新测试协议,检验模型是否准确理解文本描述的物体类别
- 在多类共存的真实图像上,发现主流模型错误率超40%且依赖干扰项
- 适合关注视觉-语言对齐、开放世界计数的开发者与研究者
开放世界文本引导的无类别计数(CAC)通过自然语言提示灵活统计任意物体类别。然而,当前评估主要关注单类图像中的计数误差,忽略了关键问题:模型能否正确将文本提示与视觉场景对齐。本文揭示,多个先进模型在识别应计数的物体类别时表现不佳,导致语义与视觉表示错位,产生虚假计数结果,降低实际应用可靠性。为此,我们提出新评估框架:(i) PrACo++(提示感知计数++),包含负标签测试和干扰物测试两套协议及专用指标;(ii) MUCCA(多类别无类别计数)数据集,收录真实场景中每幅图像含多个标注类别的图像,不同于以往单类别基准。对10个前沿方法的全面评估显示,尽管标准计数指标表现良好,模型在语义理解与对象定位上仍存在显著缺陷。最后,我们定量分析了提示间语义相似性对失败的影响。结果强调需发展更语义一致的架构,并提供未来评估的可靠基准。
原文摘要 · Abstract (English)
Open-world text-guided class-agnostic counting (CAC) has emerged as a flexible paradigm for counting arbitrary object classes by using natural language prompts. However, current evaluation protocols primarily focus on standard counting errors within single-category images, overlooking a fundamental requirement: the ability to correctly ground the textual prompt in the visual scene. In this paper, we show that several state-of-the-art CAC models often struggle to determine which object class should be counted based on the given prompt, revealing a misalignment between textual semantics and visual object representations. This limitation leads to spurious counting responses and reduced reliability in real-world scenarios. To systematically address these limitations, we propose a new evaluation framework focused on model robustness and trustworthiness. Our contribution is two-fold: (i) we introduce PrACo++ (Prompt-Aware Counting++), a novel test suite featuring two dedicated evaluation protocols -- the negative-label test and the distractor test -- paired with new specialized metrics; and (ii) we present the MUCCA (MUlti-Category Class-Agnostic counting) evaluation dataset, a new collection of real-world images featuring multiple annotated object categories per scene, unlike existing CAC benchmarks that typically include a single category per image. Our extensive experimental evaluation of 10 state-of-the-art methods shows that, despite strong performance under standard counting metrics, current models exhibit significant weaknesses in understanding and grounding object class descriptions. Finally, we provide a quantitative analysis of how semantic similarity between prompts influences these failures. Overall, our results underscore the need for more semantically grounded architectures and offer a reliable framework for future assessment in open-world text-guided CAC methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。