arXiv:2409.15953cs.CV2024-09中稿 · WACV 2025 Conferen…被引 13

新基准测试揭示提示式计数模型理解指令能力不足

Mind the Prompt: A Novel Benchmark for Prompt-based Class-Agnostic Counting

  • 构建针对性测试集与评估指标,检验模型对提示语的理解能力
  • 现有方法在标准计数指标上表现良好,但对提示理解严重偏差
  • 适合关注视觉语言模型泛化能力的研究者参考

近期物体计数研究转向无类别计数(CAC),即对训练中未见的任意物体类别进行计数。随着鲁棒视觉-语言基础模型的发展,基于提示的CAC成为热点,通过自然语言指定目标类别。然而我们发现当前基准存在显著缺陷,难以准确评估模型对提示的理解能力。主要问题在于:(i) CAC数据集多仅含单一类别图像;(ii) 现有评估器沿用传统分类特定计数方法,仅关注计数误差。为此,我们提出提示感知计数(PrACo)基准,包含两项针对性测试及专门设计的评估指标,可量化评估现有提示式CAC模型的鲁棒性与可信度。实验表明,尽管部分先进方法在标准指标上表现优异,但在理解输入提示方面存在明显缺陷,提示需更谨慎的训练或设计。代码已开源。

原文摘要 · Abstract (English)

Recently, object counting has shifted towards class-agnostic counting (CAC), which counts instances of arbitrary object classes never seen during model training. With advancements in robust vision-and-language foundation models, there is a growing interest in prompt-based CAC, where object categories are specified using natural language. However, we identify significant limitations in current benchmarks for evaluating this task, which hinder both accurate assessment and the development of more effective solutions. Specifically, we argue that the current evaluation protocols do not measure the ability of the model to understand which object has to be counted. This is due to two main factors: (i) the shortcomings of CAC datasets, which primarily consist of images containing objects from a single class, and (ii) the limitations of current counting performance evaluators, which are based on traditional class-specific counting and focus solely on counting errors. To fill this gap, we introduce the Prompt-Aware Counting (PrACo) benchmark. It comprises two targeted tests coupled with evaluation metrics specifically designed to quantitatively measure the robustness and trustworthiness of existing prompt-based CAC models. We evaluate state-of-the-art methods and demonstrate that, although some achieve impressive results on standard class-specific counting metrics, they exhibit a significant deficiency in understanding the input prompt, indicating the need for more careful training procedures or revised designs. The code for reproducing our results is available at https://github.com/ciampluca/PrACo.

计数视觉语言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。