用语音+语义结合提升课堂学生识别准确率
Multimodal Speaker Identification in Classroom Environments

- 融合语音特征与大模型生成的语义上下文进行身份锚定
- 识别准确率达50.3%,长对话可达76.9%准确率
- 可区分师生角色,适合教育智能分析场景
K-12课堂自动化分析受背景噪声和儿童语音变化影响,传统仅依赖语音的模型效果不佳。本研究评估了一种多模态说话人识别框架,将声学嵌入与大语言模型生成的语义上下文相锚定。基于EDSI数据集的8个数学课堂子集(共2,801条语句),发现仅使用声学特征的基线模型(ECAPA-TDNN)准确率为39.0%。通过在梯度提升分类器中引入基于转录文本的“上下文锚定”,多模态方法将学生识别准确率提升至50.3%。对于超过5秒的语句,准确率进一步提升至76.9%(基线为64.9%),Top-3准确率达90.9%。此外,模型对教师与学生角色的区分准确率达到99.3%。该方法推动了可考虑个体参与度的自动化反馈系统发展,是实现规模化公平教学支持的关键一步。
原文摘要 · Abstract (English)
Automated analysis of K-12 classroom dynamics faces challenges due to background noise and variable child speech, often confounding acoustic-only models. This study evaluates a multimodal speaker identification framework anchoring acoustic embeddings with LLM-derived semantic context. Using a subset of the EDSI dataset (8 math classrooms, N = 2,801 utterances), we found an acoustic baseline (ECAPA-TDNN) achieved only 39.0% accuracy. By integrating transcript-based "contextual anchoring" into a gradient boosting classifier, our multimodal approach raised student identification to 50.3%. Performance also improved for utterances over 5 seconds, reaching 76.9% accuracy (vs. 64.9% baseline) with a 90.9% Top-3 accuracy. Additionally, the model distinguished teacher vs. student roles with 99.3% accuracy. This approach advances the feasibility of automated feedback systems capable of considering individual student participation, a crucial step for supporting equitable instruction at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。