通过改写语义消除长语音中说话人风格,保护隐私
Content Anonymization for Privacy in Long-form Audio
- 用语义重写替代声纹伪装,防止多段语音关联识别
- 在电话对话场景中,内容攻击可100%还原匿名语音中的说话人
- 改写后既防识别又保语音可用性,适合会议、访谈等场景
声纹匿名技术在短句测试中表现良好,但在实际应用中,如访谈、电话和会议等长篇音频场景下,同一说话人有多个语段,攻击者可通过词汇、句法和表达习惯重新识别。为此,我们提出在ASR-TTS流程中对文本进行上下文改写,消除说话人特有风格,同时保留原意。实验表明,在长篇电话对话中,基于内容的攻击能100%复现匿名语音中的说话人身份。而所提改写方法能有效缓解此风险,且保持语音可用性。结果表明,语义改写是应对内容攻击的有效手段,建议相关方在长音频处理中引入此步骤以确保隐私安全。
原文摘要 · Abstract (English)
Voice anonymization techniques have been found to successfully obscure a speaker's acoustic identity in short, isolated utterances in benchmarks such as the VoicePrivacy Challenge. In practice, however, utterances seldom occur in isolation: long-form audio is commonplace in domains such as interviews, phone calls, and meetings. In these cases, many utterances from the same speaker are available, which pose a significantly greater privacy risk: given multiple utterances from the same speaker, an attacker could exploit an individual's vocabulary, syntax, and turns of phrase to re-identify them, even when their voice is completely disguised. To address this risk, we propose a new approach that performs a contextual rewriting of the transcripts in an ASR-TTS pipeline to eliminate speaker-specific style while preserving meaning. We present results in a long-form telephone conversation setting demonstrating the effectiveness of a content-based attack on voice-anonymized speech. Then we show how the proposed content-based anonymization methods can mitigate this risk while preserving speech utility. Overall, we find that paraphrasing is an effective defense against content-based attacks and recommend that stakeholders adopt this step to ensure anonymity in long-form audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。