arXiv:2506.16497cs.CVcs.AI2025-06

研究人脸替换视频中的视觉痕迹,发现现有模型跨数据集检测能力差

Spotting tell-tale visual artifacts in face swapping videos: strengths and pitfalls of CNN detectors

  • 用CNN模型在两个数据集上测试人脸替换的视觉痕迹检测
  • 同一数据源下检测准确率高,跨数据集时性能大幅下降
  • 提醒需针对性设计检测策略,尤其应对遮挡场景

视频中的人脸替换操作因自动化与实时工具的进步,对远程视频通信构成日益严峻的威胁。近期研究试图通过识别替换算法在复杂物理场景(如人脸遮挡)中引入的视觉痕迹来检测此类篡改。本文通过在两个数据集(含新收集的一个)上基准测试基于CNN的数据驱动模型,并分析其在不同采集来源和替换算法下的泛化能力。结果表明,通用CNN架构在相同数据源内表现优异,但在跨数据集识别遮挡相关视觉线索时存在显著困难,凸显出针对此类痕迹设计专用检测策略的必要性。

原文摘要 · Abstract (English)

Face swapping manipulations in video streams represents an increasing threat in remote video communications, due to advances in automated and real-time tools. Recent literature proposes to characterize and exploit visual artifacts introduced in video frames by swapping algorithms when dealing with challenging physical scenes, such as face occlusions. This paper investigates the effectiveness of this approach by benchmarking CNN-based data-driven models on two data corpora (including a newly collected one) and analyzing generalization capabilities with respect to different acquisition sources and swapping algorithms. The results confirm excellent performance of general-purpose CNN architectures when operating within the same data source, but a significant difficulty in robustly characterizing occlusion-based visual cues across datasets. This highlights the need for specialized detection strategies to deal with such artifacts.

人脸替换视频伪造CNN检测泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。