arXiv:2505.15004eess.AScs.SD2025-05中稿 · INTERSPEECH 2025被引 8

让语音匿名时保留情绪,防止身份泄露

EASY: Emotion-aware Speaker Anonymization via Factorized Distillation

  • 分步解耦语音的说话人身份、语言内容和情绪特征
  • 在多个数据集上同时保持高语义保真度与情绪一致性
  • 适合注重隐私与情感真实的语音匿名场景

情绪在语音交互中至关重要,通过语调、音高和节奏传递情感与意图,使交流更个性化。然而,现有语音匿名系统多采用并行解耦方法,仅分离语言内容与说话人身份,常忽视原始情绪状态的保留。本文提出EASY框架,通过新颖的顺序解耦过程,将说话人身份、语言内容和情绪表征分别建模于独立子空间,采用因子化蒸馏策略实现各属性的独立约束。该方法有效抑制信息泄漏,在语音隐私保护的同时,显著提升语言内容与情绪状态的保真度。在VoicePrivacy Challenge官方数据集上的实验表明,所提方法优于所有基线系统,兼具强隐私保护能力与高质量语音还原效果。

原文摘要 · Abstract (English)

Emotion plays a significant role in speech interaction, conveyed through tone, pitch, and rhythm, enabling the expression of feelings and intentions beyond words to create a more personalized experience. However, most existing speaker anonymization systems employ parallel disentanglement methods, which only separate speech into linguistic content and speaker identity, often neglecting the preservation of the original emotional state. In this study, we introduce EASY, an emotion-aware speaker anonymization framework. EASY employs a novel sequential disentanglement process to disentangle speaker identity, linguistic content, and emotional representation, modeling each speech attribute in distinct subspaces through a factorized distillation approach. By independently constraining speaker identity and emotional representation, EASY minimizes information leakage, enhancing privacy protection while preserving original linguistic content and emotional state. Experimental results on the VoicePrivacy Challenge official datasets demonstrate that our proposed approach outperforms all baseline systems, effectively protecting speaker privacy while maintaining linguistic content and emotional state.

语音匿名情绪保留隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。