自监督模型预训练时用荷兰语,能更好捕捉荷兰语语音特征。
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
- 在荷兰语上预训练,比英语或多语言数据更能提升语音表征
- 荷兰语语音与词汇特征在模型内部表示中更清晰可解码
- 适合关注小语种语音识别或模型本地化研究者
自监督语音模型学到的语音表征有多语言特异性?现有研究表明,仅通过语音录音训练的端到端模型可成功解码多种语言特征。但语言特定预训练对语言特异性信息编码的影响仍不明确。本文测试了自监督 Wav2Vec2 模型在荷兰语语音和词汇信息上的内部表征能力。结果表明,在纯荷兰语数据上预训练,相比等量英语或多语言数据,能显著提升对荷兰语语言特征的表征能力。该优势可通过训练好的聚类或分类探测器有效检测,部分也可通过零样本指标观察到。此外,语言特异性收益与下游自动语音识别任务性能呈正相关。
原文摘要 · Abstract (English)
How language-specific are speech representations learned by self-supervised models? Existing work has shown that a range of linguistic features can be successfully decoded from end-to-end models trained only on speech recordings. However, it's less clear to what extent pre-training on specific languages improves language-specific linguistic information. Here we test the encoding of Dutch phonetic and lexical information in internal representations of self-supervised Wav2Vec2 models. Pre-training exclusively on Dutch improves the representation of Dutch linguistic features as compared to pre-training on similar amounts of English or larger amounts of multilingual data. This language-specific advantage is well-detected by trained clustering or classification probes, and partially observable using zero-shot metrics. Furthermore, the language-specific benefit on linguistic feature encoding aligns with downstream performance on Automatic Speech Recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。