arXiv:2601.01239cs.SDcs.CR2026-01AAAI被引 3

用可逆对抗噪声保护语音隐私,防窃听又保质量

IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection

  • 用大模型生成误导性但通顺的语音内容,干扰识别
  • 针对关键词遮蔽率达96.5%(定向)和100%(非定向)
  • 恢复后音频质量高(PESQ=4.45),ASR错误率0%

人工智能快速发展推动语音识别广泛应用,但音频数据极易被未经授权访问,带来严重隐私风险。本文提出首个基于可逆对抗样本的音频隐私保护框架IO-RAE,利用大语言模型生成语义连贯但具有误导性的内容,有效防止人类与自动语音识别(ASR)系统窃听。我们引入累积信号攻击技术,通过聚焦低频信号抑制高频噪声,提升攻击效果。实验表明,该方法在多个ASR模型(包括谷歌商业黑盒系统)上实现96.5%的定向误导率和100%的非定向误导率,且恢复音频的感知语音质量(PESQ)达4.45,接近高质量原始录音;经ASR处理后错误率为0%,实现近乎无损还原。结果证明了该框架在保护敏感语音隐私方面的实用性和高效性。

原文摘要 · Abstract (English)

The rapid advancements in artificial intelligence have significantly accelerated the adoption of speech recognition technology, leading to its widespread integration across various applications. However, this surge in usage also highlights a critical issue: audio data is highly vulnerable to unauthorized exposure and analysis, posing significant privacy risks for businesses and individuals. This paper introduces an Information-Obfuscation Reversible Adversarial Example (IO-RAE) framework, the pioneering method designed to safeguard audio privacy using reversible adversarial examples. IO-RAE leverages large language models to generate misleading yet contextually coherent content, effectively preventing unauthorized eavesdropping by humans and Automatic Speech Recognition (ASR) systems. Additionally, we propose the Cumulative Signal Attack technique, which mitigates high-frequency noise and enhances attack efficacy by targeting low-frequency signals. Our approach ensures the protection of audio data without degrading its quality or our ability. Experimental evaluations demonstrate the superiority of our method, achieving a targeted misguidance rate of 96.5% and a remarkable 100% untargeted misguidance rate in obfuscating target keywords across multiple ASR models, including a commercial black-box system from Google. Furthermore, the quality of the recovered audio, measured by the Perceptual Evaluation of Speech Quality score, reached 4.45, comparable to high-quality original recordings. Notably, the recovered audio processed by ASR systems exhibited an error rate of 0%, indicating nearly lossless recovery. These results highlight the practical applicability and effectiveness of our IO-RAE framework in protecting sensitive audio privacy.

语音隐私对抗样本可逆防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。