arXiv:2507.14807cs.CVcs.AI2025-07ICCV被引 4

基于人类认知设计多人脸伪造检测框架,提升真实场景下识别准确率。

Seeing Through Deepfakes: A Human-Inspired Framework for Multi-Face Detection

  • 模仿人类观察习惯,挖掘四种关键视觉线索进行检测
  • 在基准数据集上平均准确率提升3.3%,跨数据集提升5.8%
  • 融合大模型生成可解释说明,增强结果可信度

多人脸深度伪造视频在自然社交场景中日益普遍,现有方法在单人脸检测上表现良好,但在多人脸情况下因缺乏对上下文线索的感知而失效。本文通过一系列人类实验系统研究人们在社交场景中识别深度伪造人脸的机制。定量分析揭示了四种关键线索:场景-运动一致性、人脸间外观兼容性、人际目光对齐以及人脸-身体一致性。基于这些发现,我们提出 extsf{HICOM} 框架,用于检测多人脸场景中的每一张伪造脸。在基准数据集上的实验表明, extsf{HICOM} 在同分布检测中平均准确率提升3.3%,在真实世界扰动下提升2.8%;在未见过的数据集上相比现有方法提升5.8%,体现人类启发线索的泛化能力。此外,该框架引入大语言模型(LLM)生成人类可读解释,显著提升检测结果的可解释性与说服力。本工作为融合人类因素提升深度伪造防御提供了新思路。

原文摘要 · Abstract (English)

Multi-face deepfake videos are becoming increasingly prevalent, often appearing in natural social settings that challenge existing detection methods. Most current approaches excel at single-face detection but struggle in multi-face scenarios, due to a lack of awareness of crucial contextual cues. In this work, we develop a novel approach that leverages human cognition to analyze and defend against multi-face deepfake videos. Through a series of human studies, we systematically examine how people detect deepfake faces in social settings. Our quantitative analysis reveals four key cues humans rely on: scene-motion coherence, inter-face appearance compatibility, interpersonal gaze alignment, and face-body consistency. Guided by these insights, we introduce \textsf{HICOM}, a novel framework designed to detect every fake face in multi-face scenarios. Extensive experiments on benchmark datasets show that \textsf{HICOM} improves average accuracy by 3.3\% in in-dataset detection and 2.8\% under real-world perturbations. Moreover, it outperforms existing methods by 5.8\% on unseen datasets, demonstrating the generalization of human-inspired cues. \textsf{HICOM} further enhances interpretability by incorporating an LLM to provide human-readable explanations, making detection results more transparent and convincing. Our work sheds light on involving human factors to enhance defense against deepfakes.

深度伪造多人脸检测可解释性人类认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。