arXiv:2501.05310eess.AScs.SD2025-01被引 10

分析11个自监督语音模型如何编码说话人特征,揭示深层仍保留身份信息。

A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations

  • 通过大规模探针实验分解说话人属性为声学、语调和副语言三类
  • 发现大模型深层仍能恢复说话人身份,挑战了末层仅含语言内容的共识
  • 中间表示比专用嵌入更捕捉动态语调,适合需要可解释性的任务

提升自监督语音学习(SSL)的可解释性对构建可靠语音处理系统至关重要。本研究通过大规模探针分析11个模型,将说话人身份分解为声学、语调和副语言属性。结果表明,初始层编码基础声学特征,中层合成抽象特质。关键发现是:传统认为末层仅包含纯语言内容的观点被挑战——更大模型在深层意外恢复了说话人身份。此外,语音SSL模型的中间表示比专用说话人嵌入更好地捕捉动态语调。这些发现揭示了SSL模型内部机制的复杂性,为选择可解释且任务最优的表示提供了指导。

原文摘要 · Abstract (English)

Enhancing explainability in speech self-supervised learning (SSL) is important for developing reliable SSL-based speech processing systems. This study probes how speech SSL models encode speaker-specific information via a large-scale probing analysis of 11 models, decomposing identity into acoustic, prosodic, and paralinguistic attributes. The results confirm a general hierarchy wherein initial layers encode fundamental acoustics and middle layers synthesise abstract traits. Crucially, the consensus that final layers purely abstract linguistic content is challenged. It is discovered that larger models unexpectedly recover speaker identity in their deep layers. Furthermore, the intermediate representations of speech SSL models are found to capture dynamic prosody better than specialised speaker embeddings. These insights decode the complex internal mechanics of SSL models, providing guidelines for selecting interpretable and task-optimal representations.

语音识别自监督学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。