arXiv:2505.19273cs.SDcs.AI2025-05ACL被引 3

用简单线性方程分离语音表示中的说话人信息,提升内容识别效果。

Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation

  • 通过线性分解将语音表征拆分为说话人特有与通用成分。
  • 在语音转换任务中性能超越当前最优方法。
  • 无需复杂模型,高效实现说话人身份解耦。

自监督学习(SSL)通过利用未标注数据学习有意义的语音表征,降低了语音技术对昂贵标注的依赖。由于大多数基于SSL的下游任务关注语音内容信息,理想的表征应将内容与说话人等无关变化解耦。然而,现有方法要么无法完全解耦说话人身份,要么需要资源密集型模型。本文提出一种新颖的解耦方法,通过线性方式将SSL表征分解为说话人特定和说话人无关成分,有效生成解耦表征。大量实验表明,该方法实现了良好的说话人独立性;应用于内容驱动任务如语音转换时,其表征性能显著优于当前最优方法。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has reduced the reliance on expensive labeling in speech technologies by learning meaningful representations from unannotated data. Since most SSL-based downstream tasks prioritize content information in speech, ideal representations should disentangle content from unwanted variations like speaker characteristics in the SSL representations. However, removing speaker information often degrades other speech components, and existing methods either fail to fully disentangle speaker identity or require resource-intensive models. In this paper, we propose a novel disentanglement method that linearly decomposes SSL representations into speaker-specific and speaker-independent components, effectively generating speaker disentangled representations. Comprehensive experiments show that our approach achieves speaker independence and as such, when applied to content-driven tasks such as voice conversion, our representations yield significant improvements over state-of-the-art methods.

语音处理自监督学习表征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。