arXiv:2607.21393eess.AS2026-07

用数字识别测试语音隐私防护,发现不同条件影响显著

From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers

  • 以数字识别为任务,评估三种语音混淆技术效果
  • 语音速率和攻击模型差异导致识别率相差超40%
  • 适合关注语音安全与隐私保护的研究者

在真实音频录音中保护语音隐私日益重要。本文采用数字识别作为具体且实用的评估场景,检验三种语音混淆技术在保护语言内容方面的有效性。作为基线,使用通用语音识别模型和专用数字分类器作为知情攻击者,分别识别单个数字及连续数字序列。实验结果表明,不同语音模态、语速及攻击模型下,识别性能存在显著差异。该发现凸显了构建更全面、面向应用场景的语音隐私评估方法的必要性。

原文摘要 · Abstract (English)

Protecting speech privacy in real-life audio recordings is a growing concern. This contribution evaluates the effectiveness of three obfuscation techniques in protecting linguistic speech content, using digit recognition as a task-specific and practically motivated evaluation scenario. As a first baseline, a general-purpose speech recognition model and a digit-specific classifier were applied as informed attackers to recognise both single digits and concatenated digit sequences. Our experimental results demonstrate significant differences in recognition performance across digit modality, speech rate, and attack model. These findings emphasize the need for more comprehensive and application-oriented evaluation methods to ensure speech privacy.

语音隐私数字识别对抗攻击混淆技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。