发现多语言语音识别模型存在语言数据偏见,仅依赖大语种训练权重。
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
- 用彩票理论识别模型中特定语言的子网络
- 小语种在微调时性能显著低于大语种
- 适合关注模型公平性与数据偏见的研究者
自监督学习(SSL)被广泛用于深度学习中,在无需昂贵标注的情况下训练大规模语音模型。近年来,如XLS-R这样的大型自动语音识别(ASR)模型通过同时训练超过一百种语言实现了多语言建模。然而深入分析显示,XLS-R的大部分训练数据集中于少数几种语言。尽管已有研究揭示了多领域中的偏见现象,但多语言SSL ASR中的语言偏见尚未得到充分探讨。本文采用彩票理论(LTH)识别XLS-R中与特定语言相关的子网络,并测试这些子网络在多种语言上的表现。结果表明,在微调过程中,XLS-R绕过传统语言知识,仅依赖于预训练数据中贡献最大的语言所学习到的权重。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) is used in deep learning to train on large datasets without the need for expensive labelling of the data. Recently, large Automatic Speech Recognition (ASR) models such as XLS-R have utilised SSL to train on over one hundred different languages simultaneously. However, deeper investigation shows that the bulk of the training data for XLS-R comes from a small number of languages. Biases learned through SSL have been shown to exist in multiple domains, but language bias in multilingual SSL ASR has not been thoroughly examined. In this paper, we utilise the Lottery Ticket Hypothesis (LTH) to identify language-specific subnetworks within XLS-R and test the performance of these subnetworks on a variety of different languages. We are able to show that when fine-tuning, XLS-R bypasses traditional linguistic knowledge and builds only on weights learned from the languages with the largest data contribution to the pretraining data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。