arXiv:2509.07677cs.SDcs.AI2025-09被引 2

用不可听频段操控语音,骗过语音认证与反欺骗系统

Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems

  • 在人耳不可感知的频段篡改生成语音,实现隐蔽攻击
  • 对联合认证系统攻击成功率超82%,对反欺骗系统达100%
  • 揭示现有防御体系缺陷,警示需动态演化的新一代防护

语音认证系统(VAS)利用独特的声学特征进行身份验证,正越来越多地应用于银行和医疗等高安全领域。尽管深度学习提升了其性能,但仍面临深度伪造和对抗攻击等复杂威胁。真实语音克隆使检测更加困难,现有反欺骗对策(CMs)多依赖静态模型,易被新型对抗方法绕过,形成严重安全漏洞。为此,本文提出频谱掩蔽与插值攻击(SMIA),通过精心操控人工智能生成音频中人耳不可感知的频率区域,生成听起来真实但能欺骗检测系统的对抗样本。我们在多种任务和模拟真实场景下评估该攻击对先进模型的效果:对联合语音认证/反欺骗系统攻击成功率达至少82%,对独立说话人验证系统超97.5%,对反欺骗措施则达到100%。结果明确表明当前安全机制不足以应对自适应对抗攻击,亟需转向动态、上下文感知的下一代防御框架。

原文摘要 · Abstract (English)

Voice Authentication Systems (VAS) use unique vocal characteristics for verification. They are increasingly integrated into high-security sectors such as banking and healthcare. Despite their improvements using deep learning, they face severe vulnerabilities from sophisticated threats like deepfakes and adversarial attacks. The emergence of realistic voice cloning complicates detection, as systems struggle to distinguish authentic from synthetic audio. While anti-spoofing countermeasures (CMs) exist to mitigate these risks, many rely on static detection models that can be bypassed by novel adversarial methods, leaving a critical security gap. To demonstrate this vulnerability, we propose the Spectral Masking and Interpolation Attack (SMIA), a novel method that strategically manipulates inaudible frequency regions of AI-generated audio. By altering the voice in imperceptible zones to the human ear, SMIA creates adversarial samples that sound authentic while deceiving CMs. We conducted a comprehensive evaluation of our attack against state-of-the-art (SOTA) models across multiple tasks, under simulated real-world conditions. SMIA achieved a strong attack success rate (ASR) of at least 82% against combined VAS/CM systems, at least 97.5% against standalone speaker verification systems, and 100% against countermeasures. These findings conclusively demonstrate that current security postures are insufficient against adaptive adversarial attacks. This work highlights the urgent need for a paradigm shift toward next-generation defenses that employ dynamic, context-aware frameworks capable of evolving with the threat landscape.

语音安全对抗攻击反欺骗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。