arXiv:2605.20211cs.CVcs.AI2026-05

用视觉语言模型分析学习视频中的注意力,但效果不如传统方法。

Leveraging Vision-Language Models to Detect Attention in Educational Videos

论文配图:Leveraging Vision-Language Models to Detect Attention in Educational Videos
图 1 · 摘自论文原文
  • 用VLM直接解析视频与注视数据,结合语义理解注意力焦点。
  • 在70人数据集上,所有提示策略均未超越统计基线模型。
  • 揭示了大模型在实时教育诊断中的实际局限性,适合关注可解释性的研究者。

教育视频是远程和混合学习的核心。然而,学习者的注意力波动仍是影响信息保留的关键障碍。以往研究尝试通过眼动追踪在运行时检测注意力流失,基于手工特征的机器学习分类器(如注视和扫视的统计量)进行判断。这些方法难以捕捉注意力的复杂时序特性,预测性能有限。本研究旨在通过从手工特征转向多模态基础模型,提升注意力检测能力。基于一个包含70名学习者的教育眼动数据集,我们探索了一种新方法:利用视觉语言模型(VLM)直接分析视频内容与叠加的注视数据,借助其语义推理能力,将学习者注意力置于视频语境中。我们在Gemini 3上测试多种提示策略,结果表明,所有策略均未能超越传统统计基线。该研究为VLM在实时教育诊断中的应用提供了新洞见。

原文摘要 · Abstract (English)

Educational videos are a cornerstone of remote and blended learning. However, learners' fluctuating attention remains a significant barrier to effective information retention. Prior research has attempted to mitigate this by detecting and reacting to attention loss at runtime using eye tracking. Such detection has been based so far on classical machine learning classifiers trained on engineered features, such as summary statistics over learners' fixations and saccades. These methods have struggled to capture the complex, temporal nature of learner engagement, thus exhibiting moderate prediction performance. In this study, we aim to advance the detection of attention by shifting from standard engineered features to a multimodal foundation models. Using an educational eye-tracking dataset (N = 70), we investigate a novel methodology that utilizes a Vision-Language Model (VLM) to analyze video content directly with superimposed gaze data. This approach aims to leverage the semantic reasoning capabilities of foundation models to contextualize learner focus within the video stream. We evaluate the performance of this VLM-based approach using several prompting strategies with Gemini 3, but ultimately found that none of them could outperform statistical baselines. Our results provide new insights into the limitations of using VLMs for real-time educational diagnostics.

注意力检测视觉语言模型教育技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。