arXiv:2409.00926cs.CV2024-09被引 7

新数据集+新方法,让AI更懂课堂里的学生动作。

Towards Student Actions in Classroom Scenes: New Dataset and Baseline

  • 用视觉变换器聚焦小而密的局部细节,提升动作识别精度。
  • 在4324段课堂视频上实现67.9% mAP,显著优于现有方法。
  • 适合教育智能、计算机视觉方向研究者参考。

分析学生课堂行为是教育研究中的重要且具挑战性任务。现有工作受限于缺乏可访问的数据集,难以捕捉课堂中细微的动作动态。本文提出一个全新的多标签学生动作视频(SAV)数据集,专为课堂场景中的动作检测设计。SAV数据集包含来自758个不同教室的4,324段精心剪辑的视频片段,标注了15种不同的学生动作。相比现有动作检测数据集,SAV具有真实课堂场景丰富、视频质量高、挑战性强等特点,包括细微动作差异、密集物体交互、显著尺度变化、多样拍摄角度和视觉遮挡等。这些复杂性为推进动作检测方法提供了新机遇与挑战。为此,我们提出一种基于视觉变换器的新基线方法,旨在增强对小型密集目标区域的关键局部细节的关注。该方法在SAV数据集上取得67.9%的mAP,AVA数据集上为27.4%。本文不仅发布数据集,也呼吁进一步探索人工智能驱动的教育工具,以变革教学方式与学习效果。代码与数据已开源。

原文摘要 · Abstract (English)

Analyzing student actions is an important and challenging task in educational research. Existing efforts have been hampered by the lack of accessible datasets to capture the nuanced action dynamics in classrooms. In this paper, we present a new multi-label Student Action Video (SAV) dataset, specifically designed for action detection in classroom settings. The SAV dataset consists of 4,324 carefully trimmed video clips from 758 different classrooms, annotated with 15 distinct student actions. Compared to existing action detection datasets, the SAV dataset stands out by providing a wide range of real classroom scenarios, high-quality video data, and unique challenges, including subtle movement differences, dense object engagement, significant scale differences, varied shooting angles, and visual occlusion. These complexities introduce new opportunities and challenges to advance action detection methods. To benchmark this, we propose a novel baseline method based on a visual transformer, designed to enhance attention to key local details within small and dense object regions. Our method demonstrates excellent performance with a mean Average Precision (mAP) of 67.9% and 27.4% on the SAV and AVA datasets, respectively. This paper not only provides the dataset but also calls for further research into AI-driven educational tools that may transform teaching methodologies and learning outcomes. The code and dataset are released at https://github.com/Ritatanz/SAV.

动作识别教育智能视觉变换器数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。