用视觉问答技术分析课堂行为,提升教学监控效率。
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
- 引入真实课堂视频构建VQA数据集,支持自动化行为分析
- 四种主流VQA模型在新数据集上均表现良好,准确率超预期
- 适合教育科技、智能教室研究者参考应用
课堂行为监控是教育研究的关键环节,对提升学生参与度与学习成效具有重要意义。近年来,视觉问答(VQA)模型为从视频中自动解析复杂课堂互动提供了新可能。本文探究了Llama2、Llama3、Qwen3和NVILA等前沿开源VQA模型在课堂行为分析中的适用性。为实现严谨评估,我们基于越南银行学院的真实课堂视频,构建了BAV-Classroom-VQA数据集,并详细说明了数据采集、标注方法及模型基准测试流程。初步实验结果显示,所有四个模型在回答与行为相关的视觉问题时均表现出色,展现出其在未来课堂分析与干预系统中的潜力。
原文摘要 · Abstract (English)
Classroom behavior monitoring is a critical aspect of educational research, with significant implications for student engagement and learning outcomes. Recent advancements in Visual Question Answering (VQA) models offer promising tools for automatically analyzing complex classroom interactions from video recordings. In this paper, we investigate the applicability of several state-of-the-art open-source VQA models, including LLaMA2, LLaMA3, QWEN3, and NVILA, in the context of classroom behavior analysis. To facilitate rigorous evaluation, we introduce our BAV-Classroom-VQA dataset derived from real-world classroom video recordings at the Banking Academy of Vietnam. We present the methodology for data collection, annotation, and benchmark the performance of the selected VQA models on this dataset. Our initial experimental results demonstrate that all four models achieve promising performance levels in answering behavior-related visual questions, showcasing their potential in future classroom analytics and intervention systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。