arXiv:2604.16808cs.CV2026-04

通过生物力学约束检测语音同步伪造,无需音频或像素数据。

BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling

论文配图:BioLip: Language-Generalizable Lip-Sync Deepfake Detection via Biomechanical Constraint Violation Modeling
图 1 · 摘自论文原文
  • 基于嘴唇运动的力学约束差异,捕捉时序抖动特征。
  • 零样本测试下在7种语言和5个生成器上均保持高精度。
  • 仅依赖唇部关键点,适用于跨语言、跨生成器场景。

现有语音同步深度伪造检测方法依赖像素伪影或音视频对应关系,但在生成器或语言迁移时失效,因其学习的特征受限于训练分布。本文提出新思路:真实唇动受组织力学与神经肌肉带宽约束,而当前生成模型通常不施加此类限制,导致速度、加速度和急动度方差显著高于真实语音。我们利用这一信号——称为时序唇部抖动,通过64个口周关键点在短滑动窗口内计算运动学统计量,并输入轻量三分支网络进行判断。模型仅使用关键点坐标,不依赖像素、音频或声纹数据。仅在英语数据上训练,即在零样本设置下对五个未见生成器和七种语言进行测试,展现出强泛化能力。

原文摘要 · Abstract (English)

Existing lip-sync deepfake detectors rely on pixel artifacts or audio-visual correspondence, and both fail under generator or language shift because the features they learn are tied to the training distribution. We take a different approach. Authentic lip motion is constrained by tissue mechanics and neuromuscular bandwidth; current generators typically do not impose these constraints, producing trajectories with elevated variance in velocity, acceleration, and jerk that real speech does not exhibit. We exploit this signal, which we term temporal lip jitter, by computing kinematic statistics from 64 perioral landmarks over short sliding windows and feeding them into a lightweight three-branch network. The model uses only landmark coordinates: no pixels, no audio, and no voiceprint data. We train only on English data and test in a zero-shot setting on five unseen generators and seven languages.

深度伪造检测生物力学建模跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。