用AI帮波兰语儿童筛查卷舌音发音错误,家长可操作,不替代医生诊断。
Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant
- 用wav2vec2+CTC识别语音,匹配准确率达88.7%
- 识别出错时标记替换证据,精确率72.9%,召回率61.4%
- 设计家长可用的提示助手,适合家庭早期筛查
儿童语音发音错误的早期识别常受限于专业资源不足,亟需可在临床外使用的轻量级筛查工具。本文提出针对波兰语儿童的卷舌音替代错误筛查流程,结合基于wav2vec2的CTC分词识别器、基于对齐的错误类型判定,以及基于模板的家长辅助工具,用于筛查而非诊断。在包含10名未见儿童、共559个语句的测试集上,识别器实现88.7%的完整序列匹配率。作为保守筛查指标,当系统在目标音段输出带有替换证据的标记时,对正确目标项的误报率为2.7%,精度72.9%,召回率61.4%,F1值为0.67。文中描述了助手的安全边界,并提出了未来部署中需结合临床验证的“医生在环”方案。
原文摘要 · Abstract (English)
Early identification of speech sound errors in children is often limited by access to specialists, motivating lightweight screening tools that can operate outside the clinic. We present a screening pipeline for Polish-speaking children focused on sibilant substitutions, coupling a wav2vec2-based CTC token recognizer with alignment-based error typing and a template-grounded caregiver assistant for screening, not diagnosis. On a held-out test set of 10 unseen children comprising 559 utterances, the recognizer achieves 88.7 percent exact sequence match. As a conservative screening proxy, we flag a mismatch when the system emits substitution-evidence bracketed tokens at the target segment, yielding 72.9 percent precision, 61.4 percent recall, F1 = 0.67, and a 2.7 percent false-alarm rate on target-correct items. We describe the assistant's safety boundaries and outline a clinician-in-the-loop validation plan for future deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。