让机器能按用户指定的精细程度准确数物,支持从具体到抽象的多种计数需求。
Count Anything at Any Granularity

- 将计数任务定义为多粒度问题,用文本和视觉样例共同指定目标
- 构建了迄今最大最全的计数数据集KubriCount,支持细粒度评估
- 提出HieraCount模型,在真实场景下显著提升多粒度计数准确率
开放世界物体计数仍不鲁棒:尽管视觉语言模型(VLM)进展迅速,但可靠地计数用户意图的目标物体仍远未解决。本文认为核心原因是计数粒度未显式定义——用户可能指特定身份、属性、实例类型、类别或抽象概念,但多数方法将‘计什么’视为单一类别匹配问题。为此,我们重新定义开放世界计数为多粒度计数,通过视觉样例指定目标外观,结合可选负提示的细粒度文本,明确五种语义粒度层级。然而,显式定义粒度暴露了关键数据瓶颈:现有数据集缺乏多类别场景、可控干扰项及实例级标注。为此,我们提出首个完全自动化的数据扩增流水线,融合可控3D合成、一致图像编辑与VLM过滤,构建了目前规模最大、标注最全面的计数数据集KubriCount,支持训练与多粒度评估。系统性基准测试显示,多模态大模型与专用计数模型在细粒度提示下均存在严重误执行。受此启发,我们训练了HieraCount,一种联合利用文本与视觉样例作为互补目标描述的多粒度计数模型。HieraCount显著提升多粒度计数精度,并在复杂真实场景中表现出强泛化能力。
原文摘要 · Abstract (English)
Open-world object counting remains brittle: despite rapid advances in vision-language models (VLMs), reliably counting the objects a user intends is far from solved. We argue that a central reason is that counting granularity is left implicit; users may refer to a specific identity, an attribute, an instance type, a category, or an abstract concept, yet most methods treat "what to count" as a single, category-level matching problem. In this work, we redefine open-world counting as multi-grained counting, where visual exemplars specify target appearance and fine-grained text, with optional negative prompts, specifies the intended semantic granularity across five explicit levels. Making granularity explicit, however, exposes a critical data bottleneck: existing counting datasets lack the multi-category scenes, controlled distractors, and instance-level annotations needed to verify fine-grained prompt semantics. To address this, we propose the first fully automatic data-scaling pipeline that integrates controllable 3D synthesis with consistent image editing and VLM-based filtering, and use it to construct KubriCount, the largest and most comprehensively annotated counting dataset to date, supporting both training and multi-grained evaluation. Systematic benchmarking reveals that both multimodal large language models and specialist counting models exhibit severe prompt-following failures under fine-grained distinctions. Motivated by these findings, we train HieraCount, a multi-grained counting model that jointly leverages text and visual exemplars as complementary target specifications. HieraCount substantially improves multi-grained counting accuracy and generalizes robustly to challenging real-world scenarios. The project page is available here: https://verg-avesta.github.io/KubriCount/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。