无需重训练即可在推理时消除特定人声,保护隐私
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
- 通过控制隐藏层激活抑制目标说话人特征
- 对已见和未见说话人均有效,防音色生成
- 适合需快速响应隐私请求的语音合成场景
现代零样本语音合成模型虽具高度表现力,却带来严重犯罪风险,可生成未经同意者的声线。为此,说话人遗忘旨在按需阻止特定说话人身份的生成。现有方法依赖重训练,成本高且仅限于训练集中出现的说话人。本文提出TruS,一种无需训练的说话人遗忘框架,将范式从数据删除转向推理时控制。TruS通过调整身份相关的隐藏激活,抑制目标说话人特征,同时保留韵律、情感等其他属性。实验表明,TruS能有效防止对已见和未见拒绝者说话人的语音生成,为语音合成提供可扩展的安全保障。演示与代码已公开于 http://mmai.ewha.ac.kr/trus。
原文摘要 · Abstract (English)
Modern zero-shot text-to-speech (TTS) models offer unprecedented expressivity but also pose serious crime risks, as they can synthesize voices of individuals who never consented. In this context, speaker unlearning aims to prevent the generation of specific speaker identities upon request. Existing approaches, reliant on retraining, are costly and limited to speakers seen in the training set. We present TruS, a training-free speaker unlearning framework that shifts the paradigm from data deletion to inference-time control. TruS steers identity-specific hidden activations to suppress target speakers while preserving other attributes (e.g., prosody and emotion). Experimental results show that TruS effectively prevents voice generation on both seen and unseen opt-out speakers, establishing a scalable safeguard for speech synthesis. The demo and code are available on http://mmai.ewha.ac.kr/trus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。