专为儿童语音设计的通用大模型,提升识别与筛查准确率。
KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
- 通过两阶段训练融合发音知识,增强儿童语音识别能力。
- 在四项任务中平均准确率达87%,显著优于现有模型。
- 适合儿童语言发育评估、语音治疗及教育AI开发者使用。
随着对话式和扩散型AI的快速发展,AI在教育服务中的应用日益广泛,涵盖评分、评估工具到个性化学习系统。然而,这些技术尚未充分适配儿童语音领域,现有模型多依赖清晰成人语音数据集,难以应对幼儿或存在语言障碍儿童的独特语音挑战。为此,我们提出KidSpeak——一个面向儿童语音的多任务语音增强基础模型,可执行生成与判别任务。该框架采用两阶段训练,将发音知识融入语音编码器,在四个独立任务上实现平均87%的准确率。针对人工标注成本高、现有语音对齐工具不足的问题,我们设计了灵活自动语音对齐工具FASA,利用其构建高质量训练与评估数据集。在CHILDES数据集上,FASA使噪声儿童语音对齐质量提升13.6倍,远超人工标注水平。据我们所知,KidSpeak与FASA是首个针对儿童言语与语言治疗的综合性解决方案,兼具多功能语音大模型与强大对齐工具。
原文摘要 · Abstract (English)
With the rapid advancement of conversational and diffusion-based AI, there is a growing adoption of AI in educational services, ranging from grading and assessment tools to personalized learning systems that provide targeted support for students. However, this adaptability has yet to fully extend to the domain of children's speech, where existing models often fail due to their reliance on datasets designed for clear, articulate adult speech. Children, particularly those in early developmental stages or with speech and language pathologies, present unique challenges that current AI models and datasets are ill-equipped to handle. To address this, we introduce KidSpeak, a multi-task speech-enhanced Foundation Model capable of both generative and discriminative tasks specifically tailored to children's speech patterns. Our framework employs a two-stage training process that incorporates phonetic knowledge into the speech encoder, achieving an average accuracy of 87% across four separate tasks. Furthermore, recognizing the limitations of scalable human annotation and existing speech alignment tools, we propose the Flexible and Automatic Speech Aligner (FASA) and leverage the method to construct high quality datasets for training and evaluation. This novel alignment tool significantly improves the quality of aligned children's speech from noisy data, enhancing data quality by 13.6x compared to human annotations, as demonstrated on the CHILDES dataset. To the best of our knowledge, KidSpeak and FASA represent the first comprehensive solution designed for speech and language therapy in children, offering both a multi-purpose speech LLM and a robust alignment tool.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。