用扩散模型生成骗过语音识别的假音频,同时保留原音色特征。
DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
- 基于扩散模型的语音转换,嵌入对抗性约束生成攻击音频。
- 在LibriTTS上攻击成功率显著高于基线方法。
- 生成音频保真度高,人机均难察觉异常。
作为生物识别的一种,语音识别(SID)系统的安全性至关重要。为更真实地评估SID系统的鲁棒性,本文提出DiffAttack,一种基于扩散模型的音色保留型对抗攻击方法。该方法利用扩散语音转换(DiffVC)模型生成具有目标说话人特征的伪造音频。通过在扩散逆过程引入对抗性约束,使生成过程逐步对齐目标说话人分布,从而有效误导目标模型。实验在LibriTTS数据集上表明,相比原始DiffVC及其他方法,DiffAttack显著提升攻击成功率。客观与主观评估显示,对抗性约束未损害生成语音质量,生成音频保持自然且难以察觉。
原文摘要 · Abstract (English)
Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we aim to perform more realistic attacks in SID, which are challenging for both humans and machines to detect. In this study, we propose DiffAttack, a novel timbre-reserved adversarial attack approach that exploits the capability of a diffusion-based voice conversion (DiffVC) model to generate adversarial fake audio with distinct target speaker attribution. By introducing adversarial constraints into the generative process of the diffusion-based voice conversion model, we craft fake samples that effectively mislead target models while preserving speaker-wise characteristics. Specifically, inspired by the use of randomly sampled Gaussian noise in conventional adversarial attacks and diffusion processes, we incorporate adversarial constraints into the reverse diffusion process. These constraints subtly guide the reverse diffusion process toward aligning with the target speaker distribution. Our experiments on the LibriTTS dataset indicate that DiffAttack significantly improves the attack success rate compared to vanilla DiffVC and other methods. Moreover, objective and subjective evaluations demonstrate that introducing adversarial constraints does not compromise the speech quality generated by the DiffVC model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。