arXiv:2608.03096cs.CRcs.AI2026-08KDD

测试图像级伪造检测器在视频中的表现,发现部分能超出现有视频检测方法。

FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection

论文配图:FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection
图 1 · 摘自论文原文
  • 构建9.7万视频的基准测试集,评估图像级与视频级检测器性能。
  • 最佳图像级检测器达80.16% AUC,略胜于最强视频级检测器(79.99%)。
  • 提出IV-Bridge框架,使11个图像级检测器超越主流视频检测方案。

近期视频生成模型的发展显著加剧了深度伪造威胁,但当前深度伪造视频检测基准仍不完善。特别是图像级检测器在视频领域的有效性尚未系统评估。为此,我们提出FakeI2V-Bench,一个面向挑战性场景的视频级深度伪造检测基准,重点系统评估图像级检测器在视频域的表现。该基准包含97,548个视频,涵盖最新生成模型产出的内容,覆盖更广类别。利用此数据集,我们对八种视频级检测器和十二种代表性图像级检测器进行了系统评估。实验表明,表现最佳的图像级检测器取得80.16% AUC,略优于最强视频级检测器(79.99% AUC)。进一步提出IV-Bridge通用框架,通过随机森林结合统计特征聚合帧级预测,使十一个图像级检测器超越现有视频级方法,最优变体达93.80% AUC。FakeI2V-Bench建立了严格的深度伪造视频检测基准,并提出了扩展图像级检测器至视频域的新路径,为未来研究提供新视角。代码与数据见https://github.com/CryptoAILab/FakeI2V-Bench。

原文摘要 · Abstract (English)

Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped. In particular, the effectiveness of image-level detectors in the video domain has not been systematically assessed. To fill this gap, we present FakeI2V-Bench, a benchmark for evaluating state-of-the-art video-level deepfake detectors in challenging scenarios, with a particular focus on systematically assessing the performance of image-level deepfake detectors in the video domain. FakeI2V-Bench comprises 97,548 videos, containing content generated by the latest powerful generation models and covering a broader range of categories. Using this dataset, we conduct a systematic evaluation of eight video-level detectors and twelve representative image-level detectors. Experimental results show that the best-performing image-level detector achieves an 80.16% AUC, slightly outperforming the strongest video-level detector (i.e., 79.99% AUC). Going beyond benchmarking, we present IV-Bridge, a general framework that enhances the applicability of image-level deepfake detectors to videos. IV-Bridge employs a random forest model with statistical features to aggregate frame-level predictions, allowing eleven image-level detectors to surpass state-of-the-art video-level approaches, with the best-performing variant achieving a 93.80% AUC. Overall, FakeI2V-Bench establishes a rigorous benchmark for deepfake video detection and introduces a novel pathway for extending image-level detectors to the video domain, offering new insights and directions for future research. Code and data are available at https://github.com/CryptoAILab/FakeI2V-Bench.

深度伪造视频检测图像级检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。