无需人工标注,用AI自动识别协作学习中的注视行为。
Gaze to Insight: A Scalable AI Approach for Detecting Gaze Behaviours in Face-to-Face Collaborative Learning
- 用预训练模型实现无监督的注视行为检测
- 在真实场景中达到0.829的F1分数,对同伴和电脑注视效果好
- 适合教育科技、智能教学系统研究者使用
以往研究显示,分析协作学习中的注视行为可提供有意义的教育反馈。过去几十年,机器学习方法被用于从视频中自动检测注视行为,但通常需要大量人工标注数据,且模型在不同教学配置下表现不稳定。为此,本研究提出一种可扩展的人工智能方法,利用预训练与基础模型,在无需人工标注的情况下自动检测面对面协作学习中的注视行为。该方法采用预训练YOLO11进行人员追踪,YOLOE-26结合文本提示能力识别教育相关物体,以及Gaze-LLE模型预测注视目标。结果表明,该方法在视频数据上实现了0.829的F1分数,对面向电脑和同伴的注视表现良好,其他目标则较弱;相较于传统监督学习方法,其在复杂环境中的性能更优且更稳定,展现出更强的跨配置鲁棒性。研究还探讨了该方法在真实教学环境中支持学生协作学习的潜在价值。
原文摘要 · Abstract (English)
Previous studies have illustrated the potential of analysing gaze behaviours in collaborative learning to provide educationally meaningful information for students to reflect on their learning. Over the past decades, machine learning approaches have been developed to automatically detect gaze behaviours from video data. Yet, since these approaches often require large amounts of labelled data for training, human annotation remains necessary. Additionally, researchers have questioned the cross-configuration robustness of machine learning models developed, as training datasets often fail to encompass the full range of situations encountered in educational contexts. To address these challenges, this study proposes a scalable artificial intelligence approach that leverages pretrained and foundation models to automatically detect gaze behaviours in face-to-face collaborative learning contexts without requiring human-annotated data. The approach utilises pretrained YOLO11 for person tracking, YOLOE-26 with text-prompt capability for education-related object detection, and the Gaze-LLE model for gaze target prediction. The results indicate that the proposed approach achieves an F1-score of 0.829 in detecting students' gaze behaviours from video data, with strong performance for laptop-directed gaze and peer-directed gaze, yet weaker performance for other gaze targets. Furthermore, when compared to other supervised machine learning approaches, the proposed method demonstrates superior and more stable performance in complex contexts, highlighting its better cross-configuration robustness. The implications of this approach for supporting students' collaborative learning in real-world environments are also discussed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。