融合多模态信号与个性建模,提升在线学习中注意力预测的准确性与可靠性。
Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning

- 整合视频、音频、面部动作等多源信号,结合头姿、注视等行为特征
- 在CASED数据集上表现优于随机猜测,且提供可解释的置信度评估
- 适合教育AI研发者及需高可信度决策系统的场景
从在线辅导视频中预测学生参与度极具挑战,因其涉及行为、情感与认知等多维状态。由于个体差异大且标注主观性强,该任务尤为困难。为此,我们构建了一个多模态框架,融合预训练视频、音频和图像编码器提取的隐式时空特征,以及头姿、注视、面部动作单元、情绪与小波音频特征等结构化行为信号。通过Perceiver IO潜在瓶颈整合多模态信息,并将学生与教师个性建模为可学习嵌入的变分后验,实现跨参与者的部分池化。采用证据回归与谱归一化高斯过程分类头,实现不确定性感知的预测,提升模型鲁棒性与校准度。在CASED挑战测试集上,所有参赛方法均趋近随机水平,凸显数据集难度。在此高度模糊环境下,本框架仍保持竞争力,并首次提供可靠校准的不确定性指标,证明风险量化是教育工具部署的前提。
原文摘要 · Abstract (English)
The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A reliable prediction requires bringing together different types of behavioral signals as well as expressive cues. Through our analysis of the CASED dataset, it is clear that engagement prediction gets even harder due to the high inter-person variability as well as the subjectivity of the engagement annotation. To tackle these challenges, we develop a multimodal framework that integrates the implicit spatiotemporal features extracted from pretrained video, audio, and image encoders along with structured behavioral modalities like head pose, gaze, facial action units, emotion, and wavelet-based audio features. We integrate these modalities via a Perceiver IO latent bottleneck. Moreover, student and instructor personalities are modeled as variational posteriors over learnable embeddings to enable partial pooling across participants. We employ evidential regression and spectral-normalized Gaussian process classification heads for uncertainty-aware prediction to further improve robustness and calibration. Benchmark on the CASED challenge test set shows that all participating methods converge near random-chance performance, revealing the difficulty of the dataset. In this highly ambiguous regime, our framework achieves competitive performance while uniquely offering well-calibrated uncertainty metrics, demonstrating that reliable risk-quantification is an essential prerequisite for deploying engagement models in real-world educational tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。