arXiv:2607.07322cs.CVcs.AI2026-07被引 1

为朝觐人群计数难题构建秒级标注基准,验证零样本模型实际表现。

HAJJv2-CrowdCount: Zero-Shot Benchmark for Dense Crowd Counting

  • 构建每秒人工标注的朝觐视频计数基准,解决数据缺失问题。
  • 点式计数器在最密集帧上表现最优(误差114.9),优于检测与分割方法。
  • 揭示密集遮挡场景下模型性能反转,指导实际部署决策。

朝觐视频中的人群自动计数困难并非源于模型能力不足,而是因为视频拍摄角度陡峭近乎垂直,个体间严重遮挡,单帧人数常超千人。现有测试基准或为私有、或缺乏逐秒细节。我们重新审视 HAJJv2 数据集,提出 HAJJv2-CrowdCount:对测试视频进行逐秒人工标注的人群计数。基于此,我们评估三种近期零样本计数范式:开放词汇检测器(YOLO-World)、点式计数器(APGCC)和可提示分割计数器(SAM3Count)。SAM3Count 总体均方误差最低(MAE 70.4,95% CI 56.0–86.1),优于 YOLO-World(92.0)和 APGCC(152.9)。然而,在最贴近实际部署的密集帧场景中,检测与分割类方法性能急剧下降(误差超300),而点式计数器衰减更平缓(MAE 114.9)。这一反转对朝觐人群管理至关重要,因高密度遮挡场景下需最可靠计数。标注数据已公开,以支持结果复现与拓展。

原文摘要 · Abstract (English)

Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumptions those models were built on: cameras observe the crowd from steep, near-vertical angles, individuals occlude one another extensively, and a single frame can contain well over a thousand people. Benchmarks that test crowd counting in such an environment are either private or not detailed per second. We revisit the HAJJv2 dataset and contribute HAJJv2-CrowdCount: per-second human-annotated crowd counts for its testing videos. Using these annotations, we benchmark three recent zero-shot counting paradigms: an open-vocabulary detector (YOLO-World), a point-based counter (APGCC), and a promptable segmentation-based counter (SAM3Count). SAM3Count attains the lowest overall mean absolute error (MAE 70.4, 95% CI 56.0-86.1), ahead of YOLO-World (92.0) and APGCC (152.9). This ordering reverses, however, in the regime most relevant to deployment: on the densest frames, the detection- and segmentation-based counters both degrade sharply (MAE exceeding 300), while the point-based counter degrades far more gracefully (MAE 114.9). This inversion is decision-relevant for Hajj crowd management, where reliable counts are needed most precisely in the densest and most occluded scenes. The annotations are released to support reproduction and extension of these results.

人群计数零样本视觉基准朝觐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。