用大模型知识指导提示学习,提升深伪人脸检测效果
Knowledge-Guided Prompt Learning for Deepfake Facial Image Detection
- 从大模型提取伪造相关提示作为先验知识引导优化
- 测试时微调提示,有效缓解真实场景下的领域偏移
- 在DeepFakeFaceForensics上显著超越现有方法
近期生成模型在合成照片方面表现卓越,使人类难以区分其与真实图像,尤其在逼真的人脸合成图像上。以往工作多聚焦于从海量视觉数据中挖掘判别性痕迹,但通常缺乏对先验知识的利用,且很少关注训练类别(如自然与室内物体)与测试类别(如细粒度人脸图像)之间的领域差异,导致检测性能不佳。为此,我们提出一种新型的知识引导提示学习方法用于深伪人脸检测。具体而言,我们从大型语言模型中检索与伪造相关的提示作为专家知识,指导可学习提示的优化;同时,在测试阶段进行提示微调,以缓解领域偏移问题,显著提升性能,增强实际应用能力。在DeepFakeFaceForensics数据集上的大量实验表明,所提方法明显优于当前最优方法。
原文摘要 · Abstract (English)
Recent generative models demonstrate impressive performance on synthesizing photographic images, which makes humans hardly to distinguish them from pristine ones, especially on realistic-looking synthetic facial images. Previous works mostly focus on mining discriminative artifacts from vast amount of visual data. However, they usually lack the exploration of prior knowledge and rarely pay attention to the domain shift between training categories (e.g., natural and indoor objects) and testing ones (e.g., fine-grained human facial images), resulting in unsatisfactory detection performance. To address these issues, we propose a novel knowledge-guided prompt learning method for deepfake facial image detection. Specifically, we retrieve forgery-related prompts from large language models as expert knowledge to guide the optimization of learnable prompts. Besides, we elaborate test-time prompt tuning to alleviate the domain shift, achieving significant performance improvement and boosting the application in real-world scenarios. Extensive experiments on DeepFakeFaceForensics dataset show that our proposed approach notably outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。