arXiv:2502.06460cs.CV2025-02被引 2

用不确定性建模提升复杂人群重识别准确率

Group-CLIP Uncertainty Modeling for Group Re-Identification

  • 通过模拟成员缺失和布局变化生成不确定文本描述
  • 在多个数据集上超越现有方法,最高新指标达87.3%
  • 适合关注人群结构变化的智能监控场景

群体重识别(Group ReID)旨在跨非重叠摄像头匹配行人组。与单人重识别不同,群体重识别更关注群体结构的变化,强调成员数量和空间排列。然而,多数方法依赖确定性模型,仅考虑图像中的特定群体结构,难以匹配未见的群体配置。为此,我们提出一种新型的群体CLIP不确定性建模方法(GCUM),将群体文本描述适配于不确定的成员和布局变化。具体而言,设计了成员变异模拟(MVS)模块,利用伯努利分布模拟成员缺失;群体布局自适应(GLA)模块生成带身份特异性标记的不确定群体文本描述。此外,设计群体关系构建编码器(GRCE),利用群体特征优化个体特征,并采用跨模态对比损失从群体文本描述中获取可泛化的知识。值得注意的是,我们首次将CLIP应用于群体重识别,大量实验表明,GCUM显著优于当前最先进的群体重识别方法。

原文摘要 · Abstract (English)

Group Re-Identification (Group ReID) aims matching groups of pedestrians across non-overlapping cameras. Unlike single-person ReID, Group ReID focuses more on the changes in group structure, emphasizing the number of members and their spatial arrangement. However, most methods rely on certainty-based models, which consider only the specific group structures in the group images, often failing to match unseen group configurations. To this end, we propose a novel Group-CLIP UncertaintyModeling (GCUM) approach that adapts group text descriptions to undetermined accommodate member and layout variations. Specifically, we design a Member Variant Simulation (MVS)module that simulates member exclusions using a Bernoulli distribution and a Group Layout Adaptation (GLA) module that generates uncertain group text descriptions with identity-specific tokens. In addition, we design a Group RelationshipConstruction Encoder (GRCE) that uses group features to refine individual features, and employ cross-modal contrastive loss to obtain generalizable knowledge from group text descriptions. It is worth noting that we are the first to employ CLIP to GroupReID, and extensive experiments show that GCUM significantly outperforms state-of-the-art Group ReID methods.

群体重识别不确定性建模CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。