不用说话人嵌入,用多聚类模型实现语音匿名化
SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
- 用多个针对不同说话人子集训练的聚类模型随机处理语音特征
- 相比单个聚类模型,保留更完整的语言和情感信息
- 适合关注隐私保护与攻击防御的语音系统设计者
语音匿名化在保留语言和副语言内容的同时隐藏说话人身份。自监督学习(SSL)表示虽包含语言特征,但会保留说话人特性。本文提出无需说话人嵌入的新框架SEF-MK:不使用全局训练的单一k-means模型,而是对每个语音片段随机选择一个在不同说话人子集上训练的k-means模型进行匿名化处理。从用户视角看,多聚类模型比单个模型更好保留语言与情绪内容;但从攻击者视角看,该方法反而增强了隐私攻击效果。这些发现有助于用户设计更具抗攻击能力的语音匿名系统。
原文摘要 · Abstract (English)
Voice anonymization protects speaker privacy by concealing identity while preserving linguistic and paralinguistic content. Self-supervised learning (SSL) representations encode linguistic features but preserve speaker traits. We propose a novel speaker-embedding-free framework called SEF-MK. Instead of using a single k-means model trained on the entire dataset, SEF-MK anonymizes SSL representations for each utterance by randomly selecting one of multiple k-means models, each trained on a different subset of speakers. We explore this approach from both attacker and user perspectives. Extensive experiments show that, compared to a single k-means model, SEF-MK with multiple k-means models better preserves linguistic and emotional content from the user's viewpoint. However, from the attacker's perspective, utilizing multiple k-means models boosts the effectiveness of privacy attacks. These insights can aid users in designing voice anonymization systems to mitigate attacker threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。