对比多种深度伪造检测模型,发现新方法GenConViT更准确可靠。
Comparative Analysis of Deepfake Detection Models: New Approaches and Perspectives
- 基于卷积与Transformer融合设计,提升特征捕捉能力
- 在DeepSpeak数据集上准确率达93.82%,泛化性能最优
- 适合关注虚假信息防范的媒体安全与AI研究者
深度伪造视频正严重威胁社会真实性和信息可信度,亟需高效检测技术。本文系统对比了多种深度伪造检测方法,重点分析GenConViT模型在DeepfakeBenchmark中的表现。研究涵盖深度伪造的生成原理(如CNN、GAN、Transformer)与检测技术基础,并采用WildDeepfake和DeepSpeak等新数据集进行评估。实验表明,经微调后的GenConViT在准确率(93.82%)与泛化能力方面均优于其他架构,显著提升对虚假信息传播的防御效能。本研究为构建更稳健的深度伪造检测体系提供了有效方案。
原文摘要 · Abstract (English)
The growing threat posed by deepfake videos, capable of manipulating realities and disseminating misinformation, drives the urgent need for effective detection methods. This work investigates and compares different approaches for identifying deepfakes, focusing on the GenConViT model and its performance relative to other architectures present in the DeepfakeBenchmark. To contextualize the research, the social and legal impacts of deepfakes are addressed, as well as the technical fundamentals of their creation and detection, including digital image processing, machine learning, and artificial neural networks, with emphasis on Convolutional Neural Networks (CNNs), Generative Adversarial Networks (GANs), and Transformers. The performance evaluation of the models was conducted using relevant metrics and new datasets established in the literature, such as WildDeep-fake and DeepSpeak, aiming to identify the most effective tools in the battle against misinformation and media manipulation. The obtained results indicated that GenConViT, after fine-tuning, exhibited superior performance in terms of accuracy (93.82%) and generalization capacity, surpassing other architectures in the DeepfakeBenchmark on the DeepSpeak dataset. This study contributes to the advancement of deepfake detection techniques, offering contributions to the development of more robust and effective solutions against the dissemination of false information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。