针对音视频深度伪造的细粒度检测与定位,提出高效鲁棒方案。
Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
- 融合音视频信息,设计端到端检测框架应对局部篡改。
- 在1M Deepfakes挑战赛中,定位任务排名第一,分类任务进入前四。
- 适合安全审查、内容审核等需要精准识别伪造片段的场景。
视觉与音频生成技术快速发展,新方法层出不穷,亟需稳健的合成内容检测手段。尤其当视觉、音频或两者同时被进行细粒度局部篡改时,现有检测算法面临严峻挑战。本文针对深度伪造视频的分类与定位问题提出解决方案。相关方法提交至ACM 1M Deepfakes Detection Challenge,测试集TestA的时空定位任务中取得最佳表现,分类任务排名前四。
原文摘要 · Abstract (English)
The field of visual and audio generation is burgeoning with new state-of-the-art methods. This rapid proliferation of new techniques underscores the need for robust solutions for detecting synthetic content in videos. In particular, when fine-grained alterations via localized manipulations are performed in visual, audio, or both domains, these subtle modifications add challenges to the detection algorithms. This paper presents solutions for the problems of deepfake video classification and localization. The methods were submitted to the ACM 1M Deepfakes Detection Challenge, achieving the best performance in the temporal localization task and a top four ranking in the classification task for the TestA split of the evaluation dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。