专为儿童语音设计的多语言自监督模型,提升长时录音中的说话人分割效果。
BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings
- 基于1.3万小时多语言儿童语音数据训练,适配儿童语音特征。
- 在6个语料库上取得55.0%至76.1%的F1分数,显著优于基线模型。
- 特别提升小语种表现,适合跨语言儿童语言研究者使用。
以儿童为中心的全天录音对早期语言发展研究至关重要,但现有基于成人清晰语音训练的语音模型因声学与语言差异表现不佳。本文提出BabyHuBERT,一个在40多种语言、总计13,000小时的儿童中心录音上训练的自监督语音模型。在语音类型分类任务(识别谁在说话及何时说话:主要儿童、其他儿童、男性成人、女性成人)中,BabyHuBERT-VTC在六个语料库上的F1得分介于55.0%至76.1%,持续优于W2V2-LL4300和HuBERT(分别在英语全天录音和成人语音上预训练)。在瓦努阿图和所罗门群岛的数据集上,相比HuBERT分别提升14.0和18.3绝对F1点,体现出对低资源语言的有效性。代码与模型已公开,支持跨语言儿童语音研究。
原文摘要 · Abstract (English)
Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences. We introduce BabyHuBERT, a self-supervised speech model trained on 13,000 hours of multilingual child-centered recordings from 40+ languages. Evaluated on voice type classification, the task of identifying who produces speech and when in child-centered recordings (key child, other children, male, and female adults), BabyHuBERT-VTC achieves F1-scores from 55.0% to 76.1% across six corpora, consistently outperforming W2V2-LL4300 and HuBERT (pretrained on English daylongs and clean adult speech, respectively). Notable gains include 14.0 and 18.3 absolute F1 points over HuBERT on Vanuatu and Solomon Islands, demonstrating effectiveness on underrepresented languages. We share code and models to support researchers working with child-centered recordings across diverse linguistic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。