arXiv:2602.04939cs.CV2026-02

构建首个聚焦真人表现的合成视频伪造评估基准

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes

  • 基于8个文本到视频、7个图像到视频生成器,构建2万4千余条真人视频数据集
  • 人类评估显示新基准在71%~77%对比中胜出,且面部质量指标接近真实视频
  • 现有检测模型在新基准上性能平均下降27 AUC点,适合研究伪造检测鲁棒性

现代文本到视频/图像到视频生成器已能合成极难分辨的真人视频,但现有评估体系滞后:旧基准针对篡改类伪造,近期合成视频基准则偏重规模而非真实人类表现。我们提出SynthForensics,一个以人物为中心的基准,包含来自8个T2V和7个I2V开源生成器的20,445条视频,配对来源自FF++/DFD真实视频,经两阶段人工验证,四种压缩版本并附完整元数据。在配对比较的人类实验中,评分者在71%至77%的情况下更偏好SynthForensics,优于九个现有合成视频基准;面部质量指标落在FF++/DFD基线范围内。在15种检测器和三种协议下,基于人脸的方法从FF++到SynthForensics的AUC下降13至55(均值27),进一步在强压缩下再降23;微调可缩小差距,但牺牲旧基准性能;从头训练显示多数检测器难以同时捕捉合成与篡改特征。数据集、流程和代码均已公开。

原文摘要 · Abstract (English)

Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy benchmarks target manipulation-based forgeries, and recent synthetic-video benchmarks prioritize scale over realistic human depiction. We introduce SynthForensics, a people-centric benchmark of $20{,}445$ videos from 8 T2V and 7 I2V open-source generators, paired-source from FF++/DFD reals, two-stage human-validated, in four compression versions with full metadata. In our paired-comparison human study, raters prefer SynthForensics in $71$--$77\%$ of head-to-head comparisons against each of nine existing synthetic-video benchmarks, while facial-quality metrics fall within the FF++/DFD baseline range. Across 15 detectors and three protocols, face-based methods drop $13$--$55$ AUC points (mean $27$) from FF++ to SynthForensics and a further $23$ under aggressive compression; fine-tuning closes the gap at a backward cost on legacy benchmarks; training from scratch shows synthetic and manipulation features largely disjoint for most detectors. We release dataset, pipeline, and code.

深度伪造视频生成评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。