用低干扰噪声保护语音隐私,让监听系统听不懂却人耳听不出来。
Whispering Under the Eaves: Protecting User Privacy Against Commercial and LLM-powered Automatic Speech Recognition Systems
- 在音频潜在空间生成可迁移的对抗扰动,保持音质同时骗过语音识别
- 在4个商用语音接口和多个大模型上实现90%以上误识别率,且人耳感知无异样
- 适合实时通话场景,对反制措施也有强鲁棒性,隐私保护更实用
自动语音识别(ASR)的广泛应用带来了大规模语音监控的风险,威胁用户隐私。本文提出一种新框架AudioShield,利用对抗样本抵御潜在窃听者。现有方法虽能生成对抗音频,但需离线优化,难以实时应用;而通用对抗扰动(UAP)虽可实时生成,却引入过多噪声,严重影响音质和人耳感知。AudioShield创新地在潜在空间中构建可迁移的通用对抗扰动(LS-TUAP),显著保留音频质量。同时通过目标特征适配,将目标文本特征嵌入扰动,提升跨模型转移能力。在四个商用ASR API(Google、Amazon、iFlytek、Alibaba)、三个语音助手、两个基于大模型的ASR及一个神经网络基底的ASR上评估显示,AudioShield在实现超过90%误识别率的同时,主观与客观音质评分均优于现有方法。此外,其在实时端到端场景中表现稳定,对自适应防御措施也具备强抵抗力。
原文摘要 · Abstract (English)
The widespread application of automatic speech recognition (ASR) supports large-scale voice surveillance, raising concerns about privacy among users. In this paper, we concentrate on using adversarial examples to mitigate unauthorized disclosure of speech privacy thwarted by potential eavesdroppers in speech communications. While audio adversarial examples have demonstrated the capability to mislead ASR models or evade ASR surveillance, they are typically constructed through time-intensive offline optimization, restricting their practicality in real-time voice communication. Recent work overcame this limitation by generating universal adversarial perturbations (UAPs) and enhancing their transferability for black-box scenarios. However, they introduced excessive noise that significantly degrades audio quality and affects human perception, thereby limiting their effectiveness in practical scenarios. To address this limitation and protect live users' speech against ASR systems, we propose a novel framework, AudioShield. Central to this framework is the concept of Transferable Universal Adversarial Perturbations in the Latent Space (LS-TUAP). By transferring the perturbations to the latent space, the audio quality is preserved to a large extent. Additionally, we propose target feature adaptation to enhance the transferability of UAPs by embedding target text features into the perturbations. Comprehensive evaluation on four commercial ASR APIs (Google, Amazon, iFlytek, and Alibaba), three voice assistants, two LLM-powered ASR and one NN-based ASR demonstrates the protection superiority of AudioShield over existing competitors, and both objective and subjective evaluations indicate that AudioShield significantly improves the audio quality. Moreover, AudioShield also shows high effectiveness in real-time end-to-end scenarios, and demonstrates strong resilience against adaptive countermeasures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。