语音对抗攻击通过音素混淆破坏说话人识别,揭示了语音系统脆弱性。
Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
- 从音素层面分析对抗扰动,发现其利用元音中心化和辅音替换等系统性混淆。
- 在16个发音多样的语句上,对抗音频导致识别错误并引发说话人身份漂移。
- 为语音识别与说话人验证系统提供音素感知防御的新方向,适合安全研究者参考。
语音中的对抗扰动对自动语音识别(ASR)和说话人验证构成严重威胁,通过引入人耳难以察觉的波形微调,显著改变系统输出。尽管针对端到端ASR模型的定向攻击已有广泛研究,但这些扰动的音素基础及其对说话人身份的影响仍不明确。本文以DeepSpeech为目标模型,从音素层面分析对抗音频,发现扰动利用元音中心化和辅音替换等系统性混淆,不仅导致转录错误,还削弱说话人验证所需的音素线索,引发身份漂移。在16个发音多样化的目标短语上评估结果显示,对抗音频同时造成转录错误与身份漂移,凸显了构建音素感知防御机制以保障ASR与说话人识别系统鲁棒性的必要性。
原文摘要 · Abstract (English)
Adversarial perturbations in speech pose a serious threat to automatic speech recognition (ASR) and speaker verification by introducing subtle waveform modifications that remain imperceptible to humans but can significantly alter system outputs. While targeted attacks on end-to-end ASR models have been widely studied, the phonetic basis of these perturbations and their effect on speaker identity remain underexplored. In this work, we analyze adversarial audio at the phonetic level and show that perturbations exploit systematic confusions such as vowel centralization and consonant substitutions. These distortions not only mislead transcription but also degrade phonetic cues critical for speaker verification, leading to identity drift. Using DeepSpeech as our ASR target, we generate targeted adversarial examples and evaluate their impact on speaker embeddings across genuine and impostor samples. Results across 16 phonetically diverse target phrases demonstrate that adversarial audio induces both transcription errors and identity drift, highlighting the need for phonetic-aware defenses to ensure the robustness of ASR and speaker recognition systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。