arXiv:2507.03043cs.CLcs.AI2025-07中稿 · 2026 ICASSP被引 5

用语音识别+大模型评分,精准评估儿童语言能力

K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function

  • 构建儿童语音专用的声学音素编码器与相似度模型结合框架
  • 在MyST和Multitudes数据集上分别降低10.47%和7.06%音素错误率
  • 适合儿童语言发育筛查、教育评估及可解释性要求高的场景

由于儿童音调高、发音拖长且数据有限,自动语音识别在评估幼儿语言时面临挑战。我们提出K-Function框架,结合高精度子词转录与基于大语言模型(LLM)的客观评分。其核心是儿童加权有限状态转换器(K-WFST),融合声学音素编码器与音素相似性模型,能捕捉儿童特异性发音错误,且全程可解释。K-WFST在MyST数据集上达到1.39%音素错误率,在Multitudes上为8.61%,相比贪婪搜索解码器分别提升10.47%和7.06%。高质量转录结果由LLM用于评估口语能力、发育里程碑、阅读与理解,结果与人类评估高度一致。研究证明,精确音素识别是构建有效评估体系的关键,支持儿童语言的大规模筛查。

原文摘要 · Abstract (English)

Evaluating young children's language is challenging for automatic speech recognizers due to high-pitched voices, prolonged sounds, and limited data. We introduce K-Function, a framework that combines accurate sub-word transcription with objective, Large Language Model (LLM)-driven scoring. Its core, Kids-Weighted Finite State Transducer (K-WFST), merges an acoustic phoneme encoder with a phoneme-similarity model to capture child-specific speech errors while remaining fully interpretable. K-WFST achieves a 1.39 % phoneme error rate on MyST and 8.61 % on Multitudes-an absolute improvement of 10.47 % and 7.06 % over a greedy-search decoder. These high-quality transcripts are used by an LLM to grade verbal skills, developmental milestones, reading, and comprehension, with results that align closely with human evaluators. Our findings show that precise phoneme recognition is essential for creating an effective assessment framework, enabling scalable language screening for children.

语音识别儿童语言大模型评分可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。