噪声环境下语音增强影响深度伪造检测效果,质量高未必更有效。
Investigating the Impact of Speech Enhancement on Audio Deepfake Detection in Noisy Environments
- 对比SEGAN与MetricGAN+两种增强算法对语音质量与检测性能的影响。
- MetricGAN+提升语音质量但降低检测准确率,SEGAN反其道而行之。
- 研究揭示语音增强可能引入干扰,适合关注安全防御的工程师参考。
逻辑访问(LA)攻击,又称语音深度伪造攻击,利用文本转语音(TTS)或语音转换(VC)生成伪造语音数据,严重威胁自动说话人验证(ASV)系统的安全性。本文研究语音质量与音频伪造检测系统(即LA任务)性能之间的关联。基于ASVspoof 2019挑战赛提供的LA数据集,对测试集施加不同信噪比(SNR)的噪声,训练数据保持不变。采用语音增强生成对抗网络(SEGAN)与度量优化生成对抗网络+(MetricGAN+)进行增强,通过感知语音质量(PESQ)和语音混响调制比(SRMR)评估其对语音质量的影响,并考察其对音频伪造检测性能的作用。结果表明,尽管更高语音质量通常有益于应用,但若引入不必要伪影或丢失关键信息,则可能损害下游任务。实验发现,虽然MetricGAN+在PESQ与SRMR上表现最优,但其导致的等错误率(EER)最高;而SEGAN虽语音质量最低,却实现了最低的EER,证明其更适合音频伪造检测任务。
原文摘要 · Abstract (English)
Logical Access (LA) attacks, also known as audio deepfake attacks, use Text-to-Speech (TTS) or Voice Conversion (VC) methods to generate spoofed speech data. This can represent a serious threat to Automatic Speaker Verification (ASV) systems, as intruders can use such attacks to bypass voice biometric security. In this study, we investigate the correlation between speech quality and the performance of audio spoofing detection systems (i.e., LA task). For that, the performance of two enhancement algorithms is evaluated based on two perceptual speech quality measures, namely Perceptual Evaluation of Speech Quality (PESQ) and Speech-to-Reverberation Modulation Ratio (SRMR), and in respect to their impact on the audio spoofing detection system. We adopted the LA dataset, provided in the ASVspoof 2019 Challenge, and corrupted its test set with different Signal-to-Noise Ratio (SNR) levels, while leaving the training data untouched. Enhancement was applied to attenuate the detrimental effects of noisy speech, and the performances of two models, Speech Enhancement Generative Adversarial Network (SEGAN) and Metric-Optimized Generative Adversarial Network Plus (MetricGAN+), were compared. Although we expect that speech quality will correlate well with speech applications' performance, it can also have as a side effect on downstream tasks if unwanted artifacts are introduced or relevant information is removed from the speech signal. Our results corroborate with this hypothesis, as we found that the enhancement algorithm leading to the highest speech quality scores, MetricGAN+, provided the lowest Equal Error Rate (EER) on the audio spoofing detection task, whereas the enhancement method with the lowest speech quality scores, SEGAN, led to the lowest EER, thus leading to better performance on the LA task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。