arXiv:2501.19184cs.CV2025-01中稿 · ad Elsevier's CVIU综述被引 7

首次系统梳理无需类别先验的物体计数方法,覆盖从参考样本到文本提示的全场景方案。

A Survey on Class-Agnostic Counting: Advancements from Reference-Based to Open-World Text-Guided Approaches

  • 按目标类别指定方式分为三类:基于参考、无参考和文本引导
  • 在FSC-147与CARPK数据集上评估30种架构,建立基准性能榜单
  • 适合关注开放词汇计数、多场景泛化能力的研究者

视觉物体计数正转向类别无关计数(CAC),旨在对任意类别物体进行计数,这是实现灵活、通用计数系统的关键能力。与人类无需先验知识即可识别并计数多样化物体不同,现有方法大多局限于已知类别的实例计数,依赖大量标注数据训练,在开放词汇设置下表现不佳。相比之下,CAC可在少量样本设定下计数训练中未见的类别。本文首次全面综述CAC方法,提出基于目标类别指定方式的三类范式:基于参考、无参考和开放世界文本引导。基于参考的方法通过示例引导机制达到当前最优性能;无参考方法通过图像内在模式消除示例依赖;文本引导方法利用视觉语言模型,通过文本提示描述物体类别,提供灵活且有前景的解决方案。基于该分类体系,我们概述了30种CAC架构,并报告其在标准基准上的表现,包括FSC-147数据集(使用金标准指标)和CARPK数据集(评估泛化能力)。最后,我们讨论持续挑战如标注依赖与泛化问题,并提出未来方向。

原文摘要 · Abstract (English)

Visual object counting has recently shifted towards class-agnostic counting (CAC), which addresses the challenge of counting objects across arbitrary categories, a crucial capability for flexible and generalizable counting systems. Unlike humans, who effortlessly identify and count objects from diverse categories without prior knowledge, most existing counting methods are restricted to enumerating instances of known classes, requiring extensive labeled datasets for training and struggling in open-vocabulary settings. In contrast, CAC aims to count objects belonging to classes never seen during training, operating in a few-shot setting. In this paper, we present the first comprehensive review of CAC methodologies. We propose a taxonomy to categorize CAC approaches into three paradigms based on how target object classes can be specified: reference-based, reference-less, and open-world text-guided. Reference-based approaches achieve state-of-the-art performance by relying on exemplar-guided mechanisms. Reference-less methods eliminate exemplar dependency by leveraging inherent image patterns. Finally, open-world text-guided methods use vision-language models, enabling object class descriptions via textual prompts, offering a flexible and promising solution. Based on this taxonomy, we provide an overview of 30 CAC architectures and report their performance on gold-standard benchmarks, discussing key strengths and limitations. Specifically, we present results on the FSC-147 dataset, setting a leaderboard using gold-standard metrics, and on the CARPK dataset to assess generalization capabilities. Finally, we offer a critical discussion of persistent challenges, such as annotation dependency and generalization, alongside future directions.

物体计数开放词汇视觉语言模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。