用提示词生成匿名语音,同时隐藏说话人和内容隐私。
SecureSpeech: Prompt-based Speaker and Content Protection
- 通过提示词控制生成与原说话人无关的声纹。
- 用大模型替换敏感信息,保留语义并提升隐私性。
- 适合需语音匿名的医疗、客服等隐私敏感场景。
随着语音领域身份盗用和说话人重识别问题日益严重,本文提出一种基于提示词的语音生成流程,实现说话人身份与语音内容的双重匿名化。该方法通过1)利用描述符生成与源说话人无关的声纹身份,2)结合命名实体识别模型与大语言模型,替换原始文本中的敏感内容。再将匿名后的声纹与文本输入文语转换模型,生成高保真且隐私友好的语音。实验表明,该方法在保障显著隐私保护的同时,仍保持良好的内容保留率与音频质量。本文还研究了不同说话人描述对生成语音实用性与隐私性的影响,以评估潜在偏差。
原文摘要 · Abstract (English)
Given the increasing privacy concerns from identity theft and the re-identification of speakers through content in the speech field, this paper proposes a prompt-based speech generation pipeline that ensures dual anonymization of both speaker identity and spoken content. This is addressed through 1) generating a speaker identity unlinkable to the source speaker, controlled by descriptors, and 2) replacing sensitive content within the original text using a name entity recognition model and a large language model. The pipeline utilizes the anonymized speaker identity and text to generate high-fidelity, privacy-friendly speech via a text-to-speech synthesis model. Experimental results demonstrate an achievement of significant privacy protection while maintaining a decent level of content retention and audio quality. This paper also investigates the impact of varying speaker descriptions on the utility and privacy of generated speech to determine potential biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。