arXiv:2505.09972eess.AScs.LG2025-05被引 5

用AI自动分析幼儿园课堂对话,准确识别说话人和语言特征。

Who Said What WSW 2.0? Enhanced Automated Analysis of Preschool Classroom Speech

  • 结合wav2vec2与Whisper模型,自动识别儿童与教师说话人。
  • 说话人分类准确率84.6%,转录词错误率儿童23.8%、教师11.9%。
  • 可扩展至超1500小时数据,适合大规模教育研究使用。

本文提出自动化框架WSW2.0,用于分析幼儿园课堂中的语音互动,通过融合基于wav2vec2的说话人分类与Whisper(large-v2和large-v3)语音转录技术,提升准确率与可扩展性。基于235分钟音频(12名儿童160分钟,5名教师75分钟)对比专家标注,系统在说话人分类上达到加权F1分数0.845、准确率0.846、纠错后卡帕系数0.672。转录质量中等至较高,教师词错误率(WER)为0.119,儿童为0.238。对于多种课堂语言特征,包括师生平均语句长度、词汇多样性、提问行为及回应,系统与专家转录的绝对一致性相关系数(ICC)介于0.64至0.98之间。为验证可扩展性,该框架应用于覆盖两年、超过1592小时的课堂音频数据集,展现其在真实世界应用中的稳健性。结果表明,深度学习与自然语言处理技术有望革新教育研究,提供对幼儿课堂语言关键特征的精准测量,从而支持更有效的干预策略与早期语言发展支持。

原文摘要 · Abstract (English)

This paper introduces an automated framework WSW2.0 for analyzing vocal interactions in preschool classrooms, enhancing both accuracy and scalability through the integration of wav2vec2-based speaker classification and Whisper (large-v2 and large-v3) speech transcription. A total of 235 minutes of audio recordings (160 minutes from 12 children and 75 minutes from 5 teachers), were used to compare system outputs to expert human annotations. WSW2.0 achieves a weighted F1 score of .845, accuracy of .846, and an error-corrected kappa of .672 for speaker classification (child vs. teacher). Transcription quality is moderate to high with word error rates of .119 for teachers and .238 for children. WSW2.0 exhibits relatively high absolute agreement intraclass correlations (ICC) with expert transcriptions for a range of classroom language features. These include teacher and child mean utterance length, lexical diversity, question asking, and responses to questions and other utterances, which show absolute agreement intraclass correlations between .64 and .98. To establish scalability, we apply the framework to an extensive dataset spanning two years and over 1,592 hours of classroom audio recordings, demonstrating the framework's robustness for broad real-world applications. These findings highlight the potential of deep learning and natural language processing techniques to revolutionize educational research by providing accurate measures of key features of preschool classroom speech, ultimately guiding more effective intervention strategies and supporting early childhood language development.

语音分析教育科技多模态儿童语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。