发现音频视频伪造数据集中的隐藏静音陷阱,用无监督学习提升检测鲁棒性。
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning

- 仅用真实数据训练,避免模型依赖虚假静音特征
- 移除静音后检测准确率提升15%以上
- 适合关注数据偏差与鲁棒检测的研究者
高质量数据集对机器学习系统的发展和评估至关重要,尤其在深伪检测等安全关键任务中。本文揭示,当前最广泛使用的两个音视频深伪数据集存在此前未被发现的虚假特征:起始静音。伪造视频普遍以极短静音开头,仅凭此特征即可近乎完美地区分真伪样本。因此,以往的音频或音视频模型会利用该静音特征,一旦静音被移除,性能显著下降。为规避此类不必要的人工痕迹及潜在未知偏差,我们提出转向仅使用真实数据进行无监督学习。通过自监督对齐音视频表征,有效消除对数据集特异性偏见的依赖,显著提升深伪检测的鲁棒性。
原文摘要 · Abstract (English)
Good datasets are essential for developing and benchmarking any machine learning system. Their importance is even more extreme for safety critical applications such as deepfake detection - the focus of this paper. Here we reveal that two of the most widely used audio-video deepfake datasets suffer from a previously unidentified spurious feature: the leading silence. Fake videos start with a very brief moment of silence and based on this feature alone, we can separate the real and fake samples almost perfectly. As such, previous audio-only and audio-video models exploit the presence of silence in the fake videos and consequently perform worse when the leading silence is removed. To circumvent latching on such unwanted artifact and possibly other unrevealed ones we propose a shift from supervised to unsupervised learning by training models exclusively on real data. We show that by aligning self-supervised audio-video representations we remove the risk of relying on dataset-specific biases and improve robustness in deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。