arXiv:2609.07369cs.CV2026-09

新基准DF26显示,人和现有检测器已难以分辨真实与合成视频。

DF26: We Cannot Tell Fake From Real Anymore

  • 构建包含271段真实与2420段合成视频的公开数据集
  • 人类与顶尖检测器准确率接近随机猜测(约50%)
  • 揭示当前评估方法缺陷,推动更鲁棒的评测标准

我们提出DF26,一个新型基准,用于检测由最新文本到视频及图像到视频模型生成的完全合成视频。视频涵盖单人公开演讲场景,包括直面镜头录制、官方声明和演播室访谈,共包含271段真实视频和2,420段由七种现代视频模型生成的合成视频。对DF26的研究表明,人类在识别AI生成视频上的表现,以及现有最先进深度伪造检测器的性能,均接近随机猜测水平。结果凸显了当前评估协议的局限性,并强调了需要构建能明确衡量现代生成模型分布偏移鲁棒性的基准。

原文摘要 · Abstract (English)

We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to modern generative model distribution shifts.

视频检测深度伪造基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。