arXiv:2603.07551cs.SDcs.AI2026-03

让语音模型忘记特定说话人,同时不破坏其他语音质量。

Targeted Speaker Poisoning Framework in Zero-Shot Text-to-Speech

  • 通过修改模型参数,阻止生成指定说话人的声音。
  • 可有效保护最多15个说话人隐私,100个时效果下降。
  • 为生成式语音隐私提供新思路,适合关注安全的开发者。

零样本文语转换(TTS)语音克隆带来严重隐私风险,需从训练好的TTS模型中移除特定说话人身份。传统机器遗忘方法在此场景下不足,因零样本TTS能仅凭参考提示动态重建声音。我们将其形式化为语音生成说话人投毒(SGSP)任务,即在保留其他说话人语音质量的前提下,修改模型以阻止特定身份生成。我们在1、15和100个被遗忘说话人条件下评估了推理时过滤与参数修改基线方法。性能通过实用性(词错误率WER)与隐私性(AUC和遗忘说话人相似度FSSIM)的权衡来衡量。结果表明,最多可实现15个说话人的强隐私保护,但在100个时因身份重叠增加导致可扩展性受限。本研究提出新问题与评估框架,推动生成式语音隐私发展。

原文摘要 · Abstract (English)

Zero-shot Text-to-Speech (TTS) voice cloning poses severe privacy risks, demanding the removal of specific speaker identities from trained TTS models. Conventional machine unlearning is insufficient in this context, as zero-shot TTS can dynamically reconstruct voices from just reference prompts. We formalize this task as Speech Generation Speaker Poisoning (SGSP), in which we modify trained models to prevent the generation of specific identities while preserving utility for other speakers. We evaluate inference-time filtering and parameter-modification baselines across 1, 15, and 100 forgotten speakers. Performance is assessed through the trade-off between utility (WER) and privacy, quantified using AUC and Forget Speaker Similarity (FSSIM). We achieve strong privacy for up to 15 speakers but reveal scalability limits at 100 speakers due to increased identity overlap. Our study thus introduces a novel problem and evaluation framework toward further advances in generative voice privacy.

语音生成隐私保护模型投毒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。