真实通信场景下的语音伪造检测更有效,数据构建方法是关键
On Deepfake Voice Detection -- It's All in the Presentation
- 提出新数据生成与研究框架,模拟真实通信信道
- 实验室环境下检测准确率提升39%,真实场景提升57%
- 优质数据比大模型对检测效果影响更大,适合安全与语音识别研究者
尽管生成式AI使恶意语音伪造技术大幅提升,但反伪造研究进展缓慢。本文指出,当前深度伪造语音数据集与研究方法导致系统难以在真实场景泛化。主要原因是原始伪造音频与经电话等通信信道传输后的伪造音频存在差异。为此,本文提出新的数据生成与研究框架,使反伪造系统在更真实环境中表现更优。遵循该框架,实验室环境检测准确率提升39%,真实世界基准测试提升57%。此外,研究还表明,数据质量提升对检测精度的影响,远超过使用更大规模SOTA模型带来的收益。因此,科学界应优先投入资源于全面的数据收集,而非单纯追求更大模型的训练。
原文摘要 · Abstract (English)
While the technologies empowering malicious audio deepfakes have dramatically evolved in recent years due to generative AI advances, the same cannot be said of global research into spoofing (deepfake) countermeasures. This paper highlights how current deepfake datasets and research methodologies led to systems that failed to generalize to real world application. The main reason is due to the difference between raw deepfake audio, and deepfake audio that has been presented through a communication channel, e.g. by phone. We propose a new framework for data creation and research methodology, allowing for the development of spoofing countermeasures that would be more effective in real-world scenarios. By following the guidelines outlined here we improved deepfake detection accuracy by 39% in more robust and realistic lab setups, and by 57% on a real-world benchmark. We also demonstrate how improvement in datasets would have a bigger impact on deepfake detection accuracy than the choice of larger SOTA models would over smaller models; that is, it would be more important for the scientific community to make greater investment on comprehensive data collection programs than to simply train larger models with higher computational demands.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。