arXiv:2604.21760cs.CVcs.HC2026-04

通过面部动态的可解释特征,揭示深伪视频在情感表达中的行为指纹。

Interpretable facial dynamics as behavioral and perceptual traces of deepfakes

  • 基于生物行为特征提取面部运动的低维模式与时间结构。
  • 含情感表达的深伪视频检测准确率显著更高,因情感信号被系统性破坏。
  • 模型与人类判断在情感视频上一致,但策略不同,互补性强。

深度伪造检测研究多依赖深度学习,虽在基准测试中表现优异,但难以揭示真实与伪造面部行为的区别。本研究提出一种可解释方法,基于面部动态的生物行为特征,评估计算检测策略与人类感知判断的关系。我们识别出面部运动的核心低维模式,并从中提取表征时空结构的时间特征。传统机器学习分类器在此特征上实现了显著优于随机猜测的检测效果,其关键在于伪造视频中更高阶的时间不规则性更明显。值得注意的是,含情感表达的视频检测准确性远高于无情感视频。情感效价分类分析进一步表明,深伪视频中的情感信号存在系统性退化,解释了情感动态对检测的影响差异。此外,我们还考察了模型决策与人类感知之间的关系:在情感视频上,模型与人类判断趋于一致,但在非情感视频上则出现分歧;即使输出一致,其底层检测策略也不同。结果表明,换脸式深伪视频具有可测量的行为指纹,尤其在情感表达时最为显著。模型与人类感知可能提供互补而非冗余的检测路径。

原文摘要 · Abstract (English)

Deepfake detection research has largely converged on deep learning approaches that, despite strong benchmark performance, offer limited insight into what distinguishes real from manipulated facial behavior. This study presents an interpretable alternative grounded in bio-behavioral features of facial dynamics and evaluates how computational detection strategies relate to human perceptual judgments. We identify core low-dimensional patterns of facial movement, from which temporal features characterizing spatiotemporal structure were derived. Traditional machine learning classifiers trained on these features achieved modest but significant above-chance deepfake classification, driven by higher-order temporal irregularities that were more pronounced in manipulated than real facial dynamics. Notably, detection was substantially more accurate for videos containing emotive expressions than those without. An emotional valence classification analysis further indicated that emotive signals are systematically degraded in deepfakes, explaining the differential impact of emotive dynamics on detection. Furthermore, we provide an additional and often overlooked dimension of explainability by assessing the relationship between model decisions and human perceptual detection. Model and human judgments converged for emotive but diverged for non-emotive videos, and even where outputs aligned, underlying detection strategies differed. These findings demonstrate that face-swapped deepfakes carry a measurable behavioral fingerprint, most salient during emotional expression. Additionally, model-human comparisons suggest that interpretable computational features and human perception may offer complementary rather than redundant routes to detection.

深伪检测可解释性面部动态人类感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。