从内部视角解析自监督语音模型的层间动态与功能演化
InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

- 从压缩、几何、鲁棒性三维度分析每层语音表征特性
- 发现不同训练目标导致声学压缩与流形展开模式差异
- 揭示深层语义剪枝、音素核心稳定与说话人信息波动
自监督学习(SSL)模型如Wav2Vec2、HuBERT和WavLM已成为语音与音频任务的基础。尽管表现优异,其层间动态机制仍不清晰。为此,本文提出名为InsideSSL的双阶段模型中心框架。首先,从压缩(熵)、几何(曲率)、抗扰动鲁棒性三个内在层面开展无任务依赖的逐层分析,发现不同训练目标会引发不同的声学压缩与流形展开模式。其次,引入跨层生成兼容矩阵(GCM)评估功能可迁移性,揭示了稳定的音素核心、身份信息的不稳定性以及深层语义的修剪现象。线性探测进一步将模型中心视角与下游任务关联,表明层结构决定了音素、音高和说话人编码的分布方式。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models, such as Wav2Vec2, HuBERT, and WavLM, have become foundational across a wide range of speech and audio tasks. Despite their success, understanding their internal layer-wise dynamics remains an ongoing challenge. To address this, we propose a two-part model-centric framework called InsideSSL. First, we establish a task-agnostic analysis from three intrinsic per-layer perspectives: compression (entropy), geometry (curvature), and robustness to perturbations. We show that varying training objectives induce distinct regimes of acoustic compression and manifold unfolding. Second, we introduce the cross-layer Generative Compatibility Matrix (GCM) to evaluate functional transferability, exposing stable phonetic cores, identity volatility, and deep-layer semantic pruning. In addition to these evaluations, linear probing connects the model-centric perspective to downstream tasks, demonstrating how layer topology dictates phoneme, pitch, and speaker encoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。