arXiv:2410.23955eess.AScs.LG2024-10

分析多分辨率语音自监督模型,发现性能提升主要来自低分辨率损失,而非下采样本身。

An Empirical Analysis of Speech Self-Supervised Learning at Multiple Resolutions

  • 用CCA和互信息分析多分辨率模型各层表征
  • 低分辨率损失是性能提升主因,下采样不提升下游效果
  • 适合研究语音表示学习机制与模型效率的学者

自监督学习(SSL)模型在语音处理中日益重要,近期进展聚焦于捕捉多时间尺度表示的架构。这类多尺度架构旨在利用语音的层次性:低分辨率成分试图捕获从音素到词再到句子等越来越抽象的概念。尽管多尺度方法相较于单尺度模型已展示出一定改进,但其提升原因缺乏充分的实证支持。本研究对多分辨率HuBERT(MR-HuBERT)的层间表征进行了初步分析,采用典型相关分析(CCA)和互信息(MI)。结果表明:(1)在SUPERB任务上的性能提升主要归因于辅助的低分辨率损失,而非下采样操作本身;(2)将特征下采样至低分辨率既未提升下游任务表现,也与高层语义信息(如词汇)无显著相关性,但确实提高了计算效率。这些发现挑战了关于MR-HuBERT多尺度本质的既有假设,凸显了将计算效率与学习更优表示相分离的重要性。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) models have become crucial in speech processing, with recent advancements concentrating on developing architectures that capture representations across multiple timescales. The primary goal of these multi-scale architectures is to exploit the hierarchical nature of speech, where lower-resolution components aim to capture representations that align with increasingly abstract concepts (e.g., from phones to words to sentences). Although multi-scale approaches have demonstrated some improvements over single-scale models, the precise reasons for these enhancements have poor empirical support. In this study, we present an initial analysis of layer-wise representations in multi-scale architectures, with a focus on Canonical Correlation Analysis (CCA) and Mutual Information (MI). We apply this analysis to Multi-Resolution HuBERT (MR-HuBERT) and find that (1) the improved performance on SUPERB tasks is primarily due to the auxiliary low-resolution loss rather than the downsampling itself, and (2) downsampling to lower resolutions neither improves downstream performance nor correlates with higher-level information (e.g., words), though it does improve computational efficiency. These findings challenge assumptions about the multi-scale nature of MR-HuBERT and motivate the importance of disentangling computational efficiency from learning better representations.

自监督学习语音表征多尺度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。