用面部动作单元指导视频表示,精准识别细微局部伪造
Detecting Localized Deepfake Manipulations Using Action Unit-Guided Video Representations
- 基于面部动作单元引导时空表征,融合预训练任务特征
- 在多种生成方法上实现20%准确率提升,超越现有最佳方法
- 适合研究细粒度深度伪造检测或安全防护的从业者
随着生成建模的快速发展,深度伪造技术正不断缩小真实与合成视频之间的差距,引发严重的隐私与安全问题。除了传统的换脸和重演,近期先进方法趋向于局部编辑,如细微调整眉毛、眼型或口部表情等特定面部特征。这类细粒度篡改给现有检测模型带来挑战,因其难以捕捉局部变化。据我们所知,本文首次提出专为泛化局部编辑设计的检测方法,利用由面部动作单元(Action Units)引导的时空表征。该方法通过跨注意力机制融合随机掩码与动作单元检测等预训练任务所学表征,生成能有效编码细微局部变化的嵌入表示。在多个深度伪造生成方法上的全面评估表明,本方法仅在传统FF+数据集上训练,却在检测最新带有细粒度局部编辑的伪造视频时达到新基准,准确率较当前最优方法提升20%。此外,该方法在标准数据集上表现良好,展现出对各类局部与全局伪造的强大鲁棒性与泛化能力。
原文摘要 · Abstract (English)
With rapid advancements in generative modeling, deepfake techniques are increasingly narrowing the gap between real and synthetic videos, raising serious privacy and security concerns. Beyond traditional face swapping and reenactment, an emerging trend in recent state-of-the-art deepfake generation methods involves localized edits such as subtle manipulations of specific facial features like raising eyebrows, altering eye shapes, or modifying mouth expressions. These fine-grained manipulations pose a significant challenge for existing detection models, which struggle to capture such localized variations. To the best of our knowledge, this work presents the first detection approach explicitly designed to generalize to localized edits in deepfake videos by leveraging spatiotemporal representations guided by facial action units. Our method leverages a cross-attention-based fusion of representations learned from pretext tasks like random masking and action unit detection, to create an embedding that effectively encodes subtle, localized changes. Comprehensive evaluations across multiple deepfake generation methods demonstrate that our approach, despite being trained solely on the traditional FF+ dataset, sets a new benchmark in detecting recent deepfake-generated videos with fine-grained local edits, achieving a $20\%$ improvement in accuracy over current state-of-the-art detection methods. Additionally, our method delivers competitive performance on standard datasets, highlighting its robustness and generalization across diverse types of local and global forgeries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。