通过分析视频传输中的生物特征泄漏,实时识别AI视频会议中的身份劫持。
Unmasking Puppeteers: Leveraging Biometric Leakage to Expose Impersonation in AI-Based Videoconferencing
- 从姿态表情隐变量中分离出持久身份特征,屏蔽临时动作干扰。
- 在多个生成模型上检测准确率超90%,误报率低于5%。
- 无需查看合成视频,适合实时安防与跨模型部署。
基于AI的说话人视频会议系统通过传输紧凑的姿态-表情隐变量并在接收端重合成RGB帧来降低带宽,但该隐变量可被攻击者操控,实现实时劫持他人形象。由于每帧均为合成内容,传统深度伪造检测方法完全失效。本文提出关键观察:姿态-表情隐变量中天然包含驱动身份的生物特征信息。为此,我们设计首个不依赖重建视频的生物特征泄漏防御机制——一种姿态条件下的大间隔对比编码器,能有效分离隐变量中的持久身份线索,同时消除临时姿态和表情干扰。对解耦后的嵌入向量进行简单余弦相似度测试,即可在视频渲染过程中实时标记非法身份替换。在多个说话人生成模型上的实验表明,本方法持续优于现有防御方案,具备实时性,并在分布外场景下表现出强泛化能力。
原文摘要 · Abstract (English)
AI-based talking-head videoconferencing systems reduce bandwidth by sending a compact pose-expression latent and re-synthesizing RGB at the receiver, but this latent can be puppeteered, letting an attacker hijack a victim's likeness in real time. Because every frame is synthetic, deepfake and synthetic video detectors fail outright. To address this security problem, we exploit a key observation: the pose-expression latent inherently contains biometric information of the driving identity. Therefore, we introduce the first biometric leakage defense without ever looking at the reconstructed RGB video: a pose-conditioned, large-margin contrastive encoder that isolates persistent identity cues inside the transmitted latent while cancelling transient pose and expression. A simple cosine test on this disentangled embedding flags illicit identity swaps as the video is rendered. Our experiments on multiple talking-head generation models show that our method consistently outperforms existing puppeteering defenses, operates in real-time, and shows strong generalization to out-of-distribution scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。