arXiv:2608.10497cs.CV2026-08

让人脸识别模型更像人,能看清动作和持久特征。

SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception

论文配图:SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception
图 1 · 摘自论文原文
  • 用大模型迁移语义知识,提取人的稳定特征
  • 新设计的注意力机制捕捉动作模式,无需大量视频数据
  • 在图像和视频识别上都领先,还保持人脸识别能力

尽管基础模型在跨模态的人类识别中取得显著进展,但其主要依赖静态几何特征提取,与人类感知本质不同。因此当前模型常出现‘语义盲区’,过度关注瞬时噪声,忽略稳定的软生物特征,难以捕捉时间运动特征。为此,我们提出SapiensID 2.0,一种融合语义与时间感知的人类识别框架。为解决软生物特征标注缺失问题,我们将多模态大语言模型(MLLMs)的零样本语义知识迁移到判别性嵌入空间,通过不变特质对齐(ITA)解决空间维度不匹配,提炼核心持久特征;通过瞬时噪声解耦(TND)分离衣物等干扰因素。此外,设计了运动语义注意力头(K-SAH),在时间窗口内追踪语义块,捕捉丰富运动签名,无需大规模视频数据。大量实验表明,SapiensID 2.0在基于图像和视频的人重识别、步态识别任务中达到顶尖性能,同时保持强健的人脸识别能力。

原文摘要 · Abstract (English)

While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature extraction. This approach fundamentally diverges from human perception. Consequently, current models often suffer from "semantic blindness," overfitting to transient noise while failing to leverage invariant soft biometrics, and struggle to capture temporal motion signatures. To bridge this gap, we propose SapiensID 2.0, a human recognition framework enriched with both semantic and temporal awareness. To overcome the lack of soft-biometric annotations, we transfer zero-shot semantic knowledge from Multimodal Large Language Models (MLLMs) into a discriminative embedding space. We resolve the dimensional mismatch between these spaces using Invariant Trait Alignment (ITA) to distill core persistent traits, and Transient Noise Disentanglement (TND) to decouple artifacts like clothing. Furthermore, we design a Kinematic Semantic Attention Head (K-SAH) that extends spatial attention across temporal windows. By tracking semantic patches over time, K-SAH captures rich kinematic signatures without requiring large-scale video datasets. Extensive experiments demonstrate that SapiensID 2.0 achieves state-of-the-art performance across image- and video-based person re-identification and gait recognition, while maintaining robust face recognition capabilities.

人脸识别运动识别大模型应用语义感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。