提出新计数任务,同时统计物体个体与语义组别。
Counting Beyond Instances: A Benchmark for Group-Individual Object Counting

- 设计统一框架,同时计数个体与语义组。
- 在1330张图像上,个体计数准确但组别计数表现差。
- 利用包含关系建模,提升组级计数性能。
视觉计数通常以实例级别进行,旨在估计图像中某类物体的数量。然而,现实世界的计数常涉及由多个实例组成的更高层次语义单元,如一串葡萄、一堆盘子或一双鞋。这暴露了现有计数方法的局限性:主要关注‘计什么’,却忽视了‘在哪个语义单位上计数’。为此,我们提出群-个体物体计数(Group-Individual Object Counting, GIC),要求模型在统一框架下同时计数个体物体和语义组。为支持该任务,我们构建了真实世界基准BunchCount,包含1,330张图像、89,254个个体标注和11,065个组标注。该数据集提供同一图像内的成对个体-组标注,并明确记录每组与其组成个体的包含关系。在BunchCount上的实验表明,当前先进计数模型在个体层面表现良好,但在组别计数上显著失败。为缓解语义粒度冲突,我们提出一种计数单元引导的关系计数框架,通过利用组-个体包含关系,在训练中正则化跨粒度表示。该方法显著提升组级计数性能,同时保持个体级计数能力,为超越实例的计数任务建立强基线。
原文摘要 · Abstract (English)
Visual counting is commonly formulated at the instance level, aiming to estimate how many objects of a queried category appear in an image. However, real-world counting often involves higher-level semantic units formed by multiple instances, such as a bunch of grapes, a stack of plates, or a pair of shoes. This exposes a key limitation of existing counting formulations, which mainly focus on what to count, while largely overlooking at which semantic unit to count. We introduce Group-Individual Object Counting (GIC), a new setting that requires models to count both individual objects and semantic groups within a unified framework. To support this new task, we present BunchCount, a real-world benchmark with 1,330 images, 89,254 individual annotations, and 11,065 group annotations. BunchCount provides paired individual-group annotations within the same image and explicitly records containment relations between each group and its constituent individuals. Experiments on BunchCount show that current advanced counting models perform well on individual instances but fail to count semantic groups more accurately. To mitigate semantic granularity conflict, we propose a counting-unit guided relational counting framework, which exploits group-individual containment relations to regularize cross-granularity representations during training. Our method substantially improves group-level counting while better preserving individual-level counting ability, establishing a strong baseline for counting beyond instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。