arXiv:2509.13939cs.CV2025-09被引 6

测试AI能否按用户意图精准计数,发现现有模型在细粒度场景下表现不佳。

Can Current AI Models Count What We Mean, Not What They See? A Benchmark and Systematic Evaluation

  • 构建包含681张高分辨率图像的PairTally数据集,区分形状、颜色等细微差异。
  • 模型在细粒度计数任务中准确率不足,尤其在语义相近类别间表现差。
  • 适合研究视觉理解、意图识别与多模态模型评估的学者使用。

视觉计数是基础但具挑战性的任务,尤其在复杂场景中需针对特定对象类型进行计数。尽管近期出现的无类别计数模型和大型视觉-语言模型(VLMs)在计数任务中展现出潜力,但其在执行精细、意图驱动计数方面的能力尚不明确。本文提出PairTally基准数据集,专为评估细粒度视觉计数而设计。该数据集包含681张高分辨率图像,每张图中含两类物体,要求模型根据形状、大小、颜色或语义上的细微差别进行区分与计数。数据集涵盖跨类别(不同类别)和同类别内(密切相关的子类别)两种设置,可严格评估选择性计数能力。我们对多种先进模型进行了基准测试,包括基于样本的方法、语言提示模型和大型VLMs。结果表明,尽管技术不断进步,当前模型仍难以可靠地计数用户意图所指的对象,尤其是在细粒度和视觉模糊的情况下。PairTally为诊断和改进细粒度视觉计数系统提供了新基础。

原文摘要 · Abstract (English)

Visual counting is a fundamental yet challenging task, especially when users need to count objects of a specific type in complex scenes. While recent models, including class-agnostic counting models and large vision-language models (VLMs), show promise in counting tasks, their ability to perform fine-grained, intent-driven counting remains unclear. In this paper, we introduce PairTally, a benchmark dataset specifically designed to evaluate fine-grained visual counting. Each of the 681 high-resolution images in PairTally contains two object categories, requiring models to distinguish and count based on subtle differences in shape, size, color, or semantics. The dataset includes both inter-category (distinct categories) and intra-category (closely related subcategories) settings, making it suitable for rigorous evaluation of selective counting capabilities. We benchmark a variety of state-of-the-art models, including exemplar-based methods, language-prompted models, and large VLMs. Our results show that despite recent advances, current models struggle to reliably count what users intend, especially in fine-grained and visually ambiguous cases. PairTally provides a new foundation for diagnosing and improving fine-grained visual counting systems.

视觉计数细粒度识别多模态评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。