构建真实通信场景下的音频伪造检测基准,提升系统鲁棒性。
Benchmarking Audio Deepfake Detection Robustness in Real-world Communication Scenarios
- 设计新数据集ADD-C,覆盖多种编解码器与丢包率组合
- 实测现有模型在真实场景下性能显著下降
- 提出新型数据增强策略,有效提升检测鲁棒性
现有的音频深度伪造检测(ADD)系统因真实通信场景中的音频编解码压缩和信道传输效应导致音质严重下降,难以有效泛化。为应对这一挑战,我们构建了一个严格的基准测试体系,用于评估ADD系统在这些场景下的表现。提出了ADD-C这一新测试数据集,涵盖不同音频编解码器组合及丢包率的多种通信条件。在该数据集上对三个基线ADD模型进行基准测试,结果显示其鲁棒性显著下降。为此,我们提出一种新型数据增强(DA)策略,以提升ADD系统的鲁棒性。实验结果表明,该方法显著提升了模型在ADD-C数据集上的性能。本基准可为未来构建实用且具备强泛化能力的ADD系统提供支持。
原文摘要 · Abstract (English)
Existing Audio Deepfake Detection (ADD) systems often struggle to generalise effectively due to the significantly degraded audio quality caused by audio codec compression and channel transmission effects in real-world communication scenarios. To address this challenge, we developed a rigorous benchmark to evaluate the performance of the ADD system under such scenarios. We introduced ADD-C, a new test dataset to evaluate the robustness of ADD systems under diverse communication conditions, including different combinations of audio codecs for compression and packet loss rates. Benchmarking three baseline ADD models on the ADD-C dataset demonstrated a significant decline in robustness under such conditions. A novel Data Augmentation (DA) strategy was proposed to improve the robustness of ADD systems. Experimental results demonstrated that the proposed approach significantly enhances the performance of ADD systems on the proposed ADD-C dataset. Our benchmark can assist future efforts towards building practical and robustly generalisable ADD systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。