跨模态计数迎来新范式,但评估体系滞后于技术发展。
Object Counting Across Modalities: Taxonomies, Benchmarks, Applications, and Open Challenges

- 提出五维分类框架,系统梳理跨模态计数方法
- 发现现有数据集存在语义、时空与遮挡推理缺陷
- 呼吁构建统一评估体系,避免模型仅优化特定基准
物体计数方法已从针对特定类别的密度回归,演变为基于基础模型的开放词汇计数。当前方法可依据多种视觉与文本提示进行实例枚举。尽管这一转变具有重大概念意义,但本综述指出,对通用性的宣称已超越评估基础设施的发展。多数进展指标依赖少数饱和基准,导致模型利用统计规律而非真正理解。新提出的诊断数据集揭示了语义锚定、时间身份一致性和遮挡下的空间推理等系统性失败。为此,本文提出五轴分类法(模态、机制、提示方式、监督层级、泛化设定),并用其审计显微镜、遥感、人群计数与农业等领域的文献,将普遍挑战归纳为六大结构性矛盾。据此提出组合场景理解、主动计数代理与统一多模态评估协议的发展路线图。核心主张是建立稳健的评估基础设施,以区分开放世界泛化与基准特异性优化,而非追求简单的增量工程。
原文摘要 · Abstract (English)
Object-counting methods have rapidly shifted from class-specific density regression to open-vocabulary, foundation-model-backed counters. These methods now enumerate instances from various visual and textual prompts. While this shift marks major conceptual progress, our survey argues that claims of universal generality have outpaced the evaluative infrastructure. Most progress metrics rely on a few saturated benchmarks that models exploit for statistical regularities. Newly introduced diagnostic datasets reveal systematic failures in semantic grounding, temporal identity, and spatial reasoning with occlusion. To address these failures, we introduce a five-axis taxonomy (modality, mechanism, prompting, supervision level, and generalization setting). We use this taxonomy to audit the literature across application domains, including microscopy, remote sensing, crowd counting, and agriculture. This formalizes prevailing challenges into six structural contradictions. From these, we propose a roadmap for compositional scene understanding, active counting agents, and unified multimodal evaluation protocols. The main imperative is to build a robust evaluation infrastructure to distinguish open-world generalization from benchmark-specific optimization, rather than simple incremental engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。