arXiv:2606.14297cs.CVcs.AI2026-06

用合成数据提升朝觐人群计数模型性能,解决真实数据少且难标注问题。

Pix2Pix-Hybrid: Structure-Guided Conditional Synthesis of Hajj Crowd Images with Multi-Channel Conditioning and Weak Attribute Supervision

论文配图:Pix2Pix-Hybrid: Structure-Guided Conditional Synthesis of Hajj Crowd Images with Multi-Channel Conditioning and Weak Attribute Supervision
图 1 · 摘自论文原文
  • 基于八通道条件输入的混合生成模型,融合结构与上下文信息生成逼真图像。
  • 构建1万张高分辨率合成图像数据集,使人群计数模型平均误差降低。
  • 适用于隐私敏感场景下的数据增强,尤其适合小样本人群计数任务。

针对朝觐场景中缺乏领域特定标注图像且大规模采集数据涉及隐私问题,本文提出Pix2Pix-Hybrid(P2P-H)——一种结构引导的混合条件生成对抗网络,用于朝觐人群图像合成与数据增强。P2P-H在Pix2Pix基础上采用U-Net生成器,以八通道输入联合编码边缘、灰度等结构线索及人群密度、时间等上下文属性。为捕捉密集场景细节,引入双尺度多分辨率PatchGAN判别器。训练结合对抗损失、感知损失与特征匹配损失,并采用自适应数据增强与稳定策略。模型在993帧真实朝觐视频片段(来自60个公开视频源)上训练,条件属性自动提取以减少人工标注。基于此框架构建了包含10,000张高分辨率图像的CrowdH合成数据集。实验表明,相比Pix2Pix与StyleGAN2-ADA,P2P-H在结构保持性方面表现更优,并展现出良好跨数据集迁移能力。进一步构建包含384张真实与85张精选合成图像的CrowdH-Mix-469数据集,评估五种人群计数模型在纯真实与真实+合成数据下的表现。结果显示,合成数据显著降低所有模型的平均绝对误差(MAE),对CSRNet提升最为明显。

原文摘要 · Abstract (English)

Developing accurate crowd-counting models for Hajj pilgrimage scenes remains challenging because domain-specific annotated images are scarce and data collection during large gatherings raises privacy concerns. To address these limitations, this paper proposes Pix2Pix-Hybrid (P2P-H), a hybrid conditional GAN for structure-guided Hajj crowd-image synthesis and data augmentation. P2P-H builds on Pix2Pix and employs a U-Net generator conditioned on eight input channels that jointly encode structural cues (edges and grayscale) and contextual attributes (crowd density and time of day). To capture detailed textures in dense scenes, the framework integrates two multi-scale PatchGAN discriminators operating at different resolutions. The training procedure combines adversarial, perceptual, and feature-matching objectives with adaptive data augmentation and stabilization strategies. The model was trained on 993 real Hajj frames collected from 60 publicly available video sources, with conditioning attributes derived automatically to reduce manual labeling effort. Using this framework, we constructed CrowdH, a synthetic dataset of 10,000 high-resolution Hajj crowd images. Experimental results show that P2P-H improves structure-preserving conditional synthesis quality compared with Pix2Pix and StyleGAN2-ADA baselines and shows favorable transfer to other crowd datasets. To assess downstream utility, we further constructed CrowdH-Mix-469, an annotated mixed real-synthetic dataset comprising 384 real Hajj images and 85 selected synthetic images,and evaluated five crowd-counting models under real-only and real-plus-synthetic training. The selected synthetic data reduced MAE across all five models, with the strongest gain observed for CSRNet.

图像生成数据增强人群计数合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。