arXiv:2503.02857cs.CVcs.AI2025-03被引 78

真实世界深度伪造数据集曝光,现有效模型性能大幅下降。

Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024

  • 从社交平台收集2024年真实深度伪造内容,覆盖88个网站52种语言。
  • 开源模型在新基准上检测准确率下降超45%,视频降幅达50%。
  • 适合研究者评估模型泛化能力,或为安防、媒体行业提供参考。

随着生成式AI日益逼真,鲁棒的深度伪造检测对防范欺诈与虚假信息至关重要。尽管许多深度伪造检测模型在学术数据集上表现优异,但这些数据集已过时且无法代表真实世界中的深度伪造。本文提出Deepfake-Eval-2024,一个基于2024年社交媒体及检测平台用户上传的真实深度伪造内容构建的多模态基准。该数据集包含45小时视频、56.5小时音频和1,975张图像,涵盖最新篡改技术,内容来自88个不同网站、支持52种语言。我们发现,开源前沿检测模型在本基准上的性能急剧下降:视频模型AUC下降50%,音频下降48%,图像下降45%。商业模型及在本数据集微调的模型表现优于通用开源模型,但仍不及人工专家水平。数据集已公开于https://github.com/nuriachandra/Deepfake-Eval-2024。

原文摘要 · Abstract (English)

In the age of increasingly realistic generative AI, robust deepfake detection is essential for mitigating fraud and disinformation. While many deepfake detectors report high accuracy on academic datasets, we show that these academic benchmarks are out of date and not representative of real-world deepfakes. We introduce Deepfake-Eval-2024, a new deepfake detection benchmark consisting of in-the-wild deepfakes collected from social media and deepfake detection platform users in 2024. Deepfake-Eval-2024 consists of 45 hours of videos, 56.5 hours of audio, and 1,975 images, encompassing the latest manipulation technologies. The benchmark contains diverse media content from 88 different websites in 52 different languages. We find that the performance of open-source state-of-the-art deepfake detection models drops precipitously when evaluated on Deepfake-Eval-2024, with AUC decreasing by 50% for video, 48% for audio, and 45% for image models compared to previous benchmarks. We also evaluate commercial deepfake detection models and models finetuned on Deepfake-Eval-2024, and find that they have superior performance to off-the-shelf open-source models, but do not yet reach the accuracy of deepfake forensic analysts. The dataset is available at https://github.com/nuriachandra/Deepfake-Eval-2024.

深度伪造检测基准真实场景多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。