arXiv:2503.17984cs.CVcs.AI2025-03CVPR被引 13

通过增强数据多样性和强模型设计,显著提升弱监督人群计数精度

Taste More, Taste Better: Diverse Data and Strong Model Boost Semi-Supervised Crowd Counting

  • 用背景填充增强数据多样性,保持场景真实感
  • 引入视觉状态空间模型捕捉全局上下文,提升复杂场景计数能力
  • 结合回归与抗噪分类头,缓解人工标注噪声影响

弱监督人群计数对降低密集场景高标注成本至关重要。尽管已有基于伪标签的方法,但如何有效利用未标注数据仍具挑战。本文提出TMTB框架,强调数据与模型双重优化:首先设计适合人群计数的数据增强技术,通过背景区域修复提升数据多样性,同时保持场景完整性;其次采用视觉状态空间模型(Visual State Space Model)作为主干网络,有效捕获极密集、低光照及恶劣天气下的全局上下文信息。除传统回归头外,还引入抗噪分类头,提供更鲁棒的监督信号,缓解人工标注噪声对回归任务的影响。在四个基准数据集上进行大量实验,结果表明本方法显著优于现有最先进方法。代码已公开于https://github.com/syhien/taste_more_taste_better。

原文摘要 · Abstract (English)

Semi-supervised crowd counting is crucial for addressing the high annotation costs of densely populated scenes. Although several methods based on pseudo-labeling have been proposed, it remains challenging to effectively and accurately utilize unlabeled data. In this paper, we propose a novel framework called Taste More Taste Better (TMTB), which emphasizes both data and model aspects. Firstly, we explore a data augmentation technique well-suited for the crowd counting task. By inpainting the background regions, this technique can effectively enhance data diversity while preserving the fidelity of the entire scenes. Secondly, we introduce the Visual State Space Model as backbone to capture the global context information from crowd scenes, which is crucial for extremely crowded, low-light, and adverse weather scenarios. In addition to the traditional regression head for exact prediction, we employ an Anti-Noise classification head to provide less exact but more accurate supervision, since the regression head is sensitive to noise in manual annotations. We conduct extensive experiments on four benchmark datasets and show that our method outperforms state-of-the-art methods by a large margin. Code is publicly available on https://github.com/syhien/taste_more_taste_better.

人群计数弱监督学习数据增强视觉建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。