统一人体与人脸识别,提升复杂场景下的身份判断能力
SapiensID: Foundation for Human Recognition

- 采用动态补丁生成与语义注意力机制,实现跨姿态、尺度的统一建模
- 在多个人体重识别数据集上达到顶尖性能,长短期识别均优于专用模型
- 适合需要鲁棒身份识别的现实应用,如安防监控与跨摄像头追踪
现有身份识别系统通常依赖独立的人脸与身体分析模型,在姿态、可视性与环境变化多样的真实场景中表现受限。本文提出SapiensID,一种统一模型,通过三项创新实现更强泛化能力:(i) Retina Patch(RP)——动态生成适应主体尺度的图像补丁,确保兴趣区域的一致标记;(ii) 遮蔽识别模型(MRM),可处理变长标记序列;(iii) 语义注意力头(SAH),通过关键身体部位周围特征池化学习姿态不变表示。为支持训练,构建了包含400万张图像的WebBody4M数据集,涵盖丰富姿态与尺度变化。大量实验表明,SapiensID在多个人体重识别基准上达到领先水平,无论是短期还是长期识别任务,均优于专用模型,且与专用人脸识别系统性能相当。此外,其在新提出的跨姿态-尺度重识别挑战中建立强基线,展现出对复杂现实条件的强大泛化能力。
原文摘要 · Abstract (English)
Existing human recognition systems often rely on separate, specialized models for face and body analysis, limiting their effectiveness in real-world scenarios where pose, visibility, and context vary widely. This paper introduces SapiensID, a unified model that bridges this gap, achieving robust performance across diverse settings. SapiensID introduces (i) Retina Patch (RP), a dynamic patch generation scheme that adapts to subject scale and ensures consistent tokenization of regions of interest, (ii) a masked recognition model (MRM) that learns from variable token length, and (iii) Semantic Attention Head (SAH), an module that learns pose-invariant representations by pooling features around key body parts. To facilitate training, we introduce WebBody4M, a large-scale dataset capturing diverse poses and scale variations. Extensive experiments demonstrate that SapiensID achieves state-of-the-art results on various body ReID benchmarks, outperforming specialized models in both short-term and long-term scenarios while remaining competitive with dedicated face recognition systems. Furthermore, SapiensID establishes a strong baseline for the newly introduced challenge of Cross Pose-Scale ReID, demonstrating its ability to generalize to complex, real-world conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。