针对儿童语音差异,按年龄分组优化模型适配器,提升识别准确率。
Age-Aware Adapter Tuning for Children's Speech Recognition

- 按3-12岁不同年龄分组训练专用适配器,捕捉发育阶段差异。
- 相比通用适配器,整体词错误率降低0.3%,跨年龄平均错误率降0.8%。
- 无需真实年龄标签,预测年龄路由也能接近理想效果,适合实际部署。
儿童自动语音识别(ASR)仍具挑战性,因儿童语音与成人语音差异大,且随成长阶段变化显著。尽管适配器调优可将大型预训练ASR模型适配至儿童语音,但单一共享适配器难以充分捕捉年龄相关变化。本文首次系统研究面向儿童语音的年龄感知适配器调优,聚焦3–12岁及更年长儿童语音。提出为不同年龄组分别训练专用适配器,并与统一的年龄条件式FiLM适配器对比。在真实年龄路由下,专用适配器使整体词错误率(WER)从12.6%降至12.3%,宏观平均WER从18.4%降至17.6%,且各年龄段均获提升。进一步发现,预测年龄路由表现接近真实路由,无需真实年龄标签即可达12.3%整体WER和17.8%宏观WER。相比之下,统一的FiLM适配器增益较小,表明单一适配器难以充分建模儿童语音的发育变化。
原文摘要 · Abstract (English)
Children's automatic speech recognition (ASR) remains challenging because child speech differs from adult speech and varies substantially across developmental stages. While adapter tuning provides a promising way to adapt large pretrained ASR models to children's speech, a single shared child adapter may not fully capture age-dependent variation. In this work, we present one of the first systematic studies of age-aware adapter tuning for child ASR, focusing on speech from children aged 3--12 and older years. We propose age-specialized adapters trained separately for different age groups and compare them with a unified age-conditioned FiLM adapter. With ground-truth age routing, age-specialized adapters improve over the standard shared child adapter baseline from 12.6% to 12.3% overall word error rate (WER) and from 18.4% to 17.6% macro WER, while consistently improving WER for all age groups. We further show that predicted-age routing remains close to ground-truth routing, achieving 12.3% overall WER and 17.8% macro WER without ground-truth age labels at inference. In contrast, unified FiLM conditioning provides smaller gains, indicating that a single unified adapter may be insufficient to capture developmental variation in child speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。