用视觉变压器检测儿童行为与合作参与度,准确率达97.58%
Learning in Focus: Detecting Behavioral and Collaborative Engagement Using Vision Transformers
- 基于视觉注意力机制分析儿童眼神、互动等视觉信号
- Swin Transformer在儿童参与度分类中达到97.58%准确率
- 适用于真实教育场景的自动化行为分析,适合教育科技研究者
在早期儿童教育中,准确检测协作与行为参与对促进有意义的学习体验至关重要。本文提出一种基于人工智能的方法,利用视觉变换器(ViTs)通过眼神方向、互动和同伴协作等视觉线索自动分类儿童的参与状态。基于ChildPlay gaze数据集,模型在标注视频片段上训练,用于识别行为与协作参与状态(如:参与、未参与、协作、非协作)。我们评估了六种先进变换器模型:Vision Transformer (ViT)、Data efficient Image Transformer (DeiT)、Swin Transformer、VitGaze、APVit 和 GazeTR。其中,Swin Transformer 表现最佳,分类准确率达97.58%,展现出其在建模局部与全局注意力方面的优势。结果表明,基于变换器的架构在真实教育环境中具备可扩展的自动化参与度分析潜力。
原文摘要 · Abstract (English)
In early childhood education, accurately detecting collaborative and behavioral engagement is essential to foster meaningful learning experiences. This paper presents an AI driven approach that leverages Vision Transformers (ViTs) to automatically classify children s engagement using visual cues such as gaze direction, interaction, and peer collaboration. Utilizing the ChildPlay gaze dataset, our method is trained on annotated video segments to classify behavioral and collaborative engagement states (e.g., engaged, not engaged, collaborative, not collaborative). We evaluated six state of the art transformer models: Vision Transformer (ViT), Data efficient Image Transformer (DeiT), Swin Transformer, VitGaze, APVit and GazeTR. Among these, the Swin Transformer achieved the highest classification performance with an accuracy of 97.58 percent, demonstrating its effectiveness in modeling local and global attention. Our results highlight the potential of transformer based architectures for scalable, automated engagement analysis in real world educational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。