arXiv:2512.18263eess.AScs.AI2025-12被引 4

通过声学与语义双重匹配提升儿童语音识别效果

TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition

  • 引入声学重排序,让候选样本更贴近测试语音的音色和发音
  • 在4个儿童语音数据集上,错误率比零样本降低53.3%
  • 适合缺乏标注数据的儿童语音识别场景

儿童语音识别因声学与语言变异性大、标注数据少、与成人语音差异显著而困难。语音基础模型可通过语音上下文学习(SICL)实现无微调适应新领域。但SICL效果依赖于上下文样本的选择。本文扩展已有基于检索的方法TICL,引入声学重排序步骤,提出TICL+,使候选样本在语义和声学上均与测试输入对齐。在四个儿童语音语料库上的实验表明,TICL+相比零样本性能相对词错误率降低53.3%,相比基线TICL降低37.6%,证明结合语义与声学信息对鲁棒、可扩展的儿童语音识别至关重要。

原文摘要 · Abstract (English)

Children's speech recognition remains challenging due to substantial acoustic and linguistic variability, limited labeled data, and significant differences from adult speech. Speech foundation models can address these challenges through Speech In-Context Learning (SICL), allowing adaptation to new domains without fine-tuning. However, the effectiveness of SICL depends on how in-context examples are selected. We extend an existing retrieval-based method, Text-Embedding KNN for SICL (TICL), introducing an acoustic reranking step to create TICL+. This extension prioritizes examples that are both semantically and acoustically aligned with the test input. Experiments on four children's speech corpora show that TICL+ achieves up to a 53.3% relative word error rate reduction over zero-shot performance and 37.6% over baseline TICL, highlighting the value of combining semantic and acoustic information for robust, scalable ASR in children's speech.

语音识别儿童语音上下文学习声学匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。