首个大规模跨场景人脸生成检测数据集与多模态一致性分析框架。
GLCF: A Global-Local Multimodal Coherence Analysis Framework for Talking Face Generation Detection

- 构建全局-局部多模态一致性分析框架,融合时空特征检测伪造视频。
- 在22种伪造技术、11个场景下验证,性能优于现有方法。
- 适合做深度伪造检测、多媒体安全研究的学者和工程师参考。
说话人脸生成(TFG)技术可仅用面部图像和文本生成逼真的人物说话视频,但滥用可能带来社会风险,亟需检测方法。然而该领域受限于缺乏公开数据集。本文构建了首个大规模多场景说话人脸数据集(MSTF),涵盖22种音视频伪造技术、11个生成场景及20多个语义场景,更贴近实际应用。同时提出一种基于多模态内容全局与局部一致性分析的检测框架,引入区域聚焦平滑度检测模块(RSFDM)和差异捕捉-时间帧聚合模块(DCTAM),以评估全局时序一致性;设计视觉-音频融合模块(V-AFM)从局部时间视角分析视听一致性。大量实验验证了数据集的合理性与挑战性,也表明所提方法在多种伪造场景下优于当前主流深度伪造检测模型。
原文摘要 · Abstract (English)
Talking face generation (TFG) allows for producing lifelike talking videos of any character using only facial images and accompanying text. Abuse of this technology could pose significant risks to society, creating the urgent need for research into corresponding detection methods. However, research in this field has been hindered by the lack of public datasets. In this paper, we construct the first large-scale multi-scenario talking face dataset (MSTF), which contains 22 audio and video forgery techniques, filling the gap of datasets in this field. The dataset covers 11 generation scenarios and more than 20 semantic scenarios, closer to the practical application scenario of TFG. Besides, we also propose a TFG detection framework, which leverages the analysis of both global and local coherence in the multimodal content of TFG videos. Therefore, a region-focused smoothness detection module (RSFDM) and a discrepancy capture-time frame aggregation module (DCTAM) are introduced to evaluate the global temporal coherence of TFG videos, aggregating multi-grained spatial information. Additionally, a visual-audio fusion module (V-AFM) is designed to evaluate audiovisual coherence within a localized temporal perspective. Comprehensive experiments demonstrate the reasonableness and challenges of our datasets, while also indicating the superiority of our proposed method compared to the state-of-the-art deepfake detection approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。