MauBERT用发音特征提升多语言语音表征,更鲁棒且适应新语言。
MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery
- 基于发音特征监督预训练,跨语言学习语音表征。
- 在55种语言上测试,比现有模型更少受上下文影响。
- 仅需10小时自监督微调即可适配新语言和口语。
本文提出MauBERT,是HuBERT的多语言扩展,利用发音特征实现稳健的跨语言语音表征学习。我们在55种语言上继续进行HuBERT的预训练,基于语音到发音特征的映射提供监督信号。模型从多语言数据中学习预测发音特征或音素,生成与语言无关的表示,捕捉多语言语音特性。通过全面的ABX可区分性测试,证明MauBERT产生的表征比当前最先进的多语言自监督学习模型更具上下文不变性。此外,模型仅需少量自监督微调(10小时语音)即可有效适应未见语言和非正式语料。这为自监督语音模型注入语言先验偏见提供了有效方法。
原文摘要 · Abstract (English)
This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on a phonetic-to-articulatory feature mapping in 55 languages. Our models learn from multilingual data to predict articulatory features or phones, resulting in language-independent representations that capture multilingual phonetic properties. Through comprehensive ABX discriminability testing, we show MauBERT models produce more context-invariant representations than state-of-the-art multilingual self-supervised learning models. Additionally, the models effectively adapt to unseen languages and casual speech with minimal self-supervised fine-tuning (10 hours of speech). This establishes an effective approach for instilling linguistic inductive biases in self-supervised speech models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。