用对话中的身体关键点识别身份,抗深度伪造能力强
Two-Stream Spatial-Temporal Transformer Framework for Person Identification via Natural Conversational Keypoints
- 分空间与时间双流建模关键点静态结构和动态运动
- 特征级融合后识别准确率达94.86%,显著优于单流
- 适合用于对抗深度伪造的身份验证系统
在生成式AI兴起的背景下,传统生物识别系统面临深度伪造与人脸重演等挑战。本文提出一种双流时空变换框架,利用在线对话中可见的上半身关键点(称为对话关键点)进行身份识别。采用Sapiens姿态估计算法提取133个关键点(基于COCO-WholeBody格式),涵盖面部特征、头部姿态和手部位置。框架包含空间变换器(STR)学习关键点的空间结构模式,以及时间变换器(TTR)捕捉运动序列模式。在114人的自然对话数据集上,空间流识别准确率为80.12%,时间流为63.61%。通过共享损失函数融合达到82.22%准确率,而特征级融合(拼接双流特征图)使准确率提升至94.86%。该方法联合建模静态结构与动态行为,构建更鲁棒的身份签名,相比依赖外观的传统方法更具防伪能力。
原文摘要 · Abstract (English)
In the age of AI-driven generative technologies, traditional biometric recognition systems face unprecedented challenges, particularly from sophisticated deepfake and face reenactment techniques. In this study, we propose a Two-Stream Spatial-Temporal Transformer Framework for person identification using upper body keypoints visible during online conversations, which we term conversational keypoints. Our framework processes both spatial relationships between keypoints and their temporal evolution through two specialized branches: a Spatial Transformer (STR) that learns distinctive structural patterns in keypoint configurations, and a Temporal Transformer (TTR) that captures sequential motion patterns. Using the state-of-the-art Sapiens pose estimator, we extract 133 keypoints (based on COCO-WholeBody format) representing facial features, head pose, and hand positions. The framework was evaluated on a dataset of 114 individuals engaged in natural conversations, achieving recognition accuracies of 80.12% for the spatial stream, 63.61% for the temporal stream. We then explored two fusion strategies: a shared loss function approach achieving 82.22% accuracy, and a feature-level fusion method that concatenates feature maps from both streams, significantly improving performance to 94.86%. By jointly modeling both static anatomical relationships and dynamic movement patterns, our approach learns comprehensive identity signatures that are more robust to spoofing than traditional appearance-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。