构建大规模真实混合物体计数数据集,提升模型在工业场景下的计数能力。
The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting

- 自动生成图像与精准标注,解决人工标注成本高、误差多的问题。
- 在真实数据集上模型平均绝对误差降低20.14%(FSC-147)和18.3%(PairTally)。
- 适合从事视觉计数、工业检测与数据合成的研究者使用。
物体计数是持续研究超过十年的基础视觉任务,但当前先进模型在主导现实应用(如工业质检、产品分拣)的混合物体场景中仍表现系统性失败。我们发现这一差距主要源于训练与评估数据的局限:真实数据标注成本过高且含噪声,而现有合成数据缺乏多样性和真实性。为此,我们提出MixCount数据集与基准,专为应对当前计数模型的失效模式设计。为克服数据构建与标注的高昂成本,我们开发了一套自动化生成流程,可规模化合成图像、细粒度文本描述及像素级精确计数标注,消除了以往数据集中存在的标注模糊问题。在MixCount上评估先进计数模型时,其性能在混合物体场景中出现严重下降。更重要的是,用该合成数据训练模型后,在真实世界基准上显著提升:在FSC-147上平均绝对误差(MAE)降低20.14%,在PairTally上降低18.3%。结果表明,MixCount既是基准也是训练数据,其自动化生成管道能提供近乎无限的标注数据,有效缓解计数模型长期面临的数据瓶颈。
原文摘要 · Abstract (English)
Object counting is a foundational vision task with over a decade of dedicated research, yet state-of-the-art models still fail systematically in the mixed-object setting that dominates real-world applications such as industrial inspection and product sorting. We show that this gap is strongly driven by limitations in existing training and evaluation data: real counting datasets are prohibitively expensive to annotate and suffer from labeling noise, while existing synthetic alternatives lack diversity and realism. We address this with MixCount, a dataset and benchmark for mixed-object counting designed to target the failure modes of current counting models. To overcome the high cost of constructing and labeling such data, we develop an automatic generation pipeline that synthesizes images, fine-grained textual descriptions, and pixel-perfect counting annotations at scale, eliminating the labeling ambiguity that plagues prior datasets. Evaluating state-of-the-art counting models on MixCount exposes severe degradation in the mixed-object setting. More importantly, training these models on our synthesized data yields substantial gains on real-world benchmarks, reducing MAE by 20.14% on FSC-147 and by 18.3% on PairTally. These results establish MixCount as both a benchmark and a training dataset for fine-grained counting, and demonstrate that our pipeline, which produces effectively unlimited labeled data, helps address a long-standing bottleneck in counting models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。