用AI自动分析课堂视频与文本,实现教学活动与对话的精准识别。
Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
- 分模态设计视频与文本分析流程,分别处理24类教学活动和19类对话类型。
- 微调模型在视频和文本上分别达到0.577和0.460的宏F1分数。
- 适用于教育技术研究者及智能教学反馈系统开发者。
课堂互动观察可为教师提供具体反馈,但现有方法依赖人工标注,成本高且难以扩展。本文探索基于AI的课堂录像分析,聚焦多模态教学活动与对话识别,作为可操作反馈的基础。利用包含164小时视频和68节课程转录文本的密集标注数据集,设计并行的、模态专用分析管道:视频侧评估零样本多模态大模型、微调视觉语言模型与自监督视频变换器,覆盖24类教学活动;文本侧采用上下文感知的Transformer分类器进行微调,并与提示驱动的大模型对比,识别19类对话类型。针对类别不平衡与多标签复杂性问题,引入逐标签阈值、上下文窗口与不平衡损失函数。结果表明,微调模型始终优于提示方法,在视频和文本任务上分别取得0.577与0.460的宏F1分数。实验验证了自动化课堂分析的可行性,为可扩展的教师反馈系统奠定了基础。
原文摘要 · Abstract (English)
Observation of classroom interactions can provide concrete feedback to teachers, but current methods rely on manual annotation, which is resource-intensive and hard to scale. This work explores AI-driven analysis of classroom recordings, focusing on multimodal instructional activity and discourse recognition as a foundation for actionable feedback. Using a densely annotated dataset of 164 hours of video and 68 lesson transcripts, we design parallel, modality-specific pipelines. For video, we evaluate zero-shot multimodal LLMs, fine-tuned vision-language models, and self-supervised video transformers on 24 activity labels. For transcripts, we fine-tune a transformer-based classifier with contextualized inputs and compare it against prompting-based LLMs on 19 discourse labels. To handle class imbalance and multi-label complexity, we apply per-label thresholding, context windows, and imbalance-aware loss functions. The results show that fine-tuned models consistently outperform prompting-based approaches, achieving macro-F1 scores of 0.577 for video and 0.460 for transcripts. These results demonstrate the feasibility of automated classroom analysis and establish a foundation for scalable teacher feedback systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。