arXiv:2511.23098eess.AS2025-11被引 3

按孩子语音相似性分组,高效适配成人模型提升儿童语音识别

Group-Aware Partial Model Merging for Children's Automatic Speech Recognition

  • 先聚类孩子语音数据,再对每组部分微调并合并模型
  • 在MyST数据集上相对误识率降低6%,参数量更少
  • 适合资源有限但需精准识别儿童语音的场景

尽管使用成人预训练模型进行监督微调在儿童语音识别中展现出潜力,但通常难以捕捉儿童群体特有的特征与差异。为此,我们提出一种参数高效的组感知局部模型融合方法(GRAPAM),结合无监督聚类、局部微调和模型合并。该方法首先根据声学相似性对儿童语音数据进行分组,每组用于部分微调一个成人预训练模型,并在参数层面合并所得模型。在MyST儿童语音语料库上的实验表明,使用相同数据量时,GRAPAM实现了相对词错误率(WER)6%的提升,优于全量微调,且训练参数更少。

原文摘要 · Abstract (English)

While supervised fine-tuning of adult pre-trained models for children's ASR has shown promise, it often fails to capture group-specific characteristics and variations among children. To address this, we introduce GRoup-Aware PARtial model Merging, a parameter-efficient approach that combines unsupervised clustering, partial fine-tuning, and model merging. Our approach adapts adult-pre-trained models to children by first grouping the children's data based on acoustic similarity. Each group is used to partially fine-tune an adult pre-trained model, and the resulting models are merged at the parameter level. Experiments conducted on the MyST children's speech corpus indicate that GRAPAM achieves a relative WER improvement of 6%, using the same amount of data, outperforming full fine-tuning while training fewer parameters.

语音识别儿童语音模型融合参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。