探究自监督语音模型注意力机制,发现跨语言存在显著差异。
Probing self-attention in self-supervised speech models for cross-linguistic differences
- 分析小型自监督语音模型的注意力头,发现其分布从局部到全局不等。
- 跨语言对比显示土耳其语与英语在注意力模式上存在明显差异。
- 模型主要依赖对角注意力头进行音素分类,适用于语音表征研究。
由于新型Transformer架构的出现,语音模型在自动语音识别(ASR)基准上的性能显著提升。尽管如此,关于注意力机制在语音任务中的作用仍知之甚少。例如,尽管普遍认为这些模型学习的是语言无关(即通用)的语音表征,但对其‘语言无关’的具体含义尚未深入探讨。本文以小型自监督语音Transformer模型TERA为例,研究其自注意力机制。结果表明,即使模型规模小,其学习到的注意力头也表现出多样性,从几乎完全对角到几乎完全全局,且不受训练语言影响。我们还发现了土耳其语与英语在注意力模式上的显著差异,并证实模型在预训练过程中确实学到了重要的语音学信息。此外,通过注意力头消融实验,发现跨语言模型主要依赖对角注意力头进行音素分类。
原文摘要 · Abstract (English)
Speech models have gained traction thanks to increase in accuracy from novel transformer architectures. While this impressive increase in performance across automatic speech recognition (ASR) benchmarks is noteworthy, there is still much that is unknown about the use of attention mechanisms for speech-related tasks. For example, while it is assumed that these models are learning language-independent (i.e., universal) speech representations, there has not yet been an in-depth exploration of what it would mean for the models to be language-independent. In the current paper, we explore this question within the realm of self-attention mechanisms of one small self-supervised speech transformer model (TERA). We find that even with a small model, the attention heads learned are diverse ranging from almost entirely diagonal to almost entirely global regardless of the training language. We highlight some notable differences in attention patterns between Turkish and English and demonstrate that the models do learn important phonological information during pretraining. We also present a head ablation study which shows that models across languages primarily rely on diagonal heads to classify phonemes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。