arXiv:2501.01877cs.CV2025-01

用照片估算人群总体积,助力公共安全与基础设施评估

ANTHROPOS-V: benchmarking the novel task of Crowd Volume Estimation

  • 基于合成视频数据集,仅凭RGB图像估计人群总体积
  • 模型在真实图像上表现良好,体积估计误差可控
  • 适合关注人群密度与重量分布的智慧城市应用

我们提出全新任务——人群体积估计(Crowd Volume Estimation, CVE),即仅通过RGB图像估算人群集体身体体积。该任务不仅服务于活动管理与公共安全,还可用于估算人体重,支持基础设施应力评估与重量均衡保障等场景。为此,我们构建首个CVE基准:ANTHROPOS-V,一个包含多样化城市环境的合成逼真视频数据集,标注了每个人体的体积、SMPL形状参数与关键点。我们还探索了相关评估指标,定义了基于人体网格恢复与人群计数的基线模型,并提出一种专为CVE设计的方法,优于基线。尽管数据为合成,个体身高体重分布符合真实人口统计,且在真实图像上的下游任务中表现良好。代码与数据集已开源。

原文摘要 · Abstract (English)

We introduce the novel task of Crowd Volume Estimation (CVE), defined as the process of estimating the collective body volume of crowds using only RGB images. Besides event management and public safety, CVE can be instrumental in approximating body weight, unlocking weight sensitive applications such as infrastructure stress assessment, and assuring even weight balance. We propose the first benchmark for CVE, comprising ANTHROPOS-V, a synthetic photorealistic video dataset featuring crowds in diverse urban environments. Its annotations include each person's volume, SMPL shape parameters, and keypoints. Also, we explore metrics pertinent to CVE, define baseline models adapted from Human Mesh Recovery and Crowd Counting domains, and propose a CVE specific methodology that surpasses baselines. Although synthetic, the weights and heights of individuals are aligned with the real-world population distribution across genders, and they transfer to the downstream task of CVE from real images. Benchmark and code are available at github.com/colloroneluca/Crowd-Volume-Estimation.

人群估计体积估计合成数据城市安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。