arXiv:2507.17851cs.SDeess.AS2025-07

用可解释性量化并清除语音模型中的说话人信息残留,提升内容识别准确率与隐私安全。

Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability

  • 基于SHAP分析构建可解释性基准,直接量化内容嵌入中残留的说话人信息比例。
  • 通过噪声过滤方法将说话人信息残留从18.05%降至接近零,且语音识别损失增加不足1%。
  • 无需重训练、兼容多种模型,适合需隐私保护的语音内容处理场景。

自监督语音模型学习的表示同时包含内容与说话人信息,这种耦合导致内容任务受说话人偏差影响,且匿名化表示仍可能泄露身份。本文提出两个贡献:首先,构建基于可解释性分析的InterpTRQE-SptME基准,利用SHAP方法直接量化内容嵌入中残留的说话人信息比例;其次,提出InterpTF-SptME方法,根据该分析结果过滤嵌入中的说话人信息。在VCTK数据集上测试包括HuBERT、WavLM和ContentVec在内的七种模型,结果显示,使用SHAP噪声过滤后,说话人信息残留由18.05%降至近乎为零,同时语音识别的CTC损失增加小于1%。该方法不依赖特定模型,无需重新训练。

原文摘要 · Abstract (English)

Self-supervised speech models learn representations that capture both content and speaker information. Yet this entanglement creates problems: content tasks suffer from speaker bias, and privacy concerns arise when speaker identity leaks through supposedly anonymized representations. We present two contributions to address these challenges. First, we develop InterpTRQE-SptME (Timbre Residual Quantitative Evaluation Benchmark of Speech pre-training Models Encoding via Interpretability), a benchmark that directly measures residual speaker information in content embeddings using SHAP-based interpretability analysis. Unlike existing indirect metrics, our approach quantifies the exact proportion of speaker information remaining after disentanglement. Second, we propose InterpTF-SptME, which uses these interpretability insights to filter speaker information from embeddings. Testing on VCTK with seven models including HuBERT, WavLM, and ContentVec, we find that SHAP Noise filtering reduces speaker residuals from 18.05% to nearly zero while maintaining recognition accuracy (CTC loss increase under 1%). The method is model-agnostic and requires no retraining.

语音模型可解释性隐私保护去耦合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。