arXiv:2607.06392cs.SD2026-07中稿 · INTERSPEECH 2026被引 1

从内部视角解析自监督语音模型的层间动态与功能演化

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

论文配图:InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective
图 1 · 摘自论文原文
  • 从压缩、几何、鲁棒性三维度分析每层语音表征特性
  • 发现不同训练目标导致声学压缩与流形展开模式差异
  • 揭示深层语义剪枝、音素核心稳定与说话人信息波动

自监督学习(SSL)模型如Wav2Vec2、HuBERT和WavLM已成为语音与音频任务的基础。尽管表现优异,其层间动态机制仍不清晰。为此,本文提出名为InsideSSL的双阶段模型中心框架。首先,从压缩(熵)、几何(曲率)、抗扰动鲁棒性三个内在层面开展无任务依赖的逐层分析,发现不同训练目标会引发不同的声学压缩与流形展开模式。其次,引入跨层生成兼容矩阵(GCM)评估功能可迁移性,揭示了稳定的音素核心、身份信息的不稳定性以及深层语义的修剪现象。线性探测进一步将模型中心视角与下游任务关联,表明层结构决定了音素、音高和说话人编码的分布方式。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) models, such as Wav2Vec2, HuBERT, and WavLM, have become foundational across a wide range of speech and audio tasks. Despite their success, understanding their internal layer-wise dynamics remains an ongoing challenge. To address this, we propose a two-part model-centric framework called InsideSSL. First, we establish a task-agnostic analysis from three intrinsic per-layer perspectives: compression (entropy), geometry (curvature), and robustness to perturbations. We show that varying training objectives induce distinct regimes of acoustic compression and manifold unfolding. Second, we introduce the cross-layer Generative Compatibility Matrix (GCM) to evaluate functional transferability, exposing stable phonetic cores, identity volatility, and deep-layer semantic pruning. In addition to these evaluations, linear probing connects the model-centric perspective to downstream tasks, demonstrating how layer topology dictates phoneme, pitch, and speaker encoding.

自监督学习语音表示模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。