改进Vision Transformer模型,精准识别深度伪造图像。
Combating Digitally Altered Images: Deepfake Detection
- 基于改进ViT构建检测模型,融合多种数据增强提升鲁棒性。
- 在OpenForensics子集上实现领先准确率,有效区分真实与伪造图像。
- 适合图像安全、媒体审核领域研究人员参考使用。
深度伪造技术生成的超逼真篡改图像和视频对公众及相关部门构成重大挑战。本研究提出一种基于改进视觉变换器(Vision Transformer, ViT)的深度伪造检测方法,通过在OpenForensics数据集子集上训练,结合多种数据增强技术以提升对多样化图像篡改的鲁棒性。针对类别不平衡问题,采用过采样并进行分层抽样的训练-验证集划分。性能评估基于训练集和测试集上的准确率,并对随机人物图像的预测得分进行独立测试。实验结果表明,该模型在测试集上达到当前最优检测效果,能精细识别深度伪造图像。
原文摘要 · Abstract (English)
The rise of Deepfake technology to generate hyper-realistic manipulated images and videos poses a significant challenge to the public and relevant authorities. This study presents a robust Deepfake detection based on a modified Vision Transformer(ViT) model, trained to distinguish between real and Deepfake images. The model has been trained on a subset of the OpenForensics Dataset with multiple augmentation techniques to increase robustness for diverse image manipulations. The class imbalance issues are handled by oversampling and a train-validation split of the dataset in a stratified manner. Performance is evaluated using the accuracy metric on the training and testing datasets, followed by a prediction score on a random image of people, irrespective of their realness. The model demonstrates state-of-the-art results on the test dataset to meticulously detect Deepfake images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。