arXiv:2604.04482cs.AI2026-04中稿 · the 27th Internati…

用多模态大模型预测学习者视频行为,实现可解释的教育设计预评估。

Scalable and Explainable Learner-Video Interaction Prediction using Multimodal Large Language Models

论文配图:Scalable and Explainable Learner-Video Interaction Prediction using Multimodal Large Language Models
图 1 · 摘自论文原文
  • 基于多模态大模型提取视频片段嵌入,识别观看/暂停等行为高峰。
  • 在66门课程7700万条操作数据上验证,模型能准确预测行为峰值并跨学科泛化。
  • 通过概念激活向量解释预测结果,契合多媒体学习理论,适合教育研究者使用。

学习者在教育视频中使用播放控制的行为可反映认知负荷与教学设计质量,但现有预测模型缺乏可扩展性和可解释性,限制了教师部署前的预判能力。本文提出一种可扩展、可解释的流水线,仅基于视频内容预测群体层面的观看、暂停、跳过和倒放行为,作为认知负荷的代理指标。方法利用多模态大语言模型(MLLM)对短视频段生成嵌入,并训练神经分类器识别时间精细的行为高峰。结合多媒体学习理论中的教学设计原则,使用GPT-5为视频片段编码特征,并通过概念激活向量(CAVs)解释模型预测。在来自66门在线课程的7700万条视频控制事件上进行评估,结果表明:基于MLLM嵌入的分类器能可靠预测行为高峰,具备跨学术领域的泛化能力,并编码出与理论相关的可解释教学概念。整体表明,该方法实现了低成本、可解释的教育视频设计预筛选,为大规模实证检验多媒体学习理论开辟新路径。

原文摘要 · Abstract (English)

Learners' use of video controls in educational videos provides implicit signals of cognitive processing and instructional design quality, yet the lack of scalable and explainable predictive models limits instructors' ability to anticipate such behavior before deployment. We propose a scalable, interpretable pipeline for predicting population-level watching, pausing, skipping, and rewinding behavior as proxies for cognitive load from video content alone. Our approach leverages multimodal large language models (MLLMs) to compute embeddings of short video segments and trains a neural classifier to identify temporally fine-grained interaction peaks. Drawing from multimedia learning theory on instructional design for optimal cognitive load, we code features of the video segments using GPT-5 and employ them as a basis for interpreting model predictions via concept activation vectors. We evaluate our pipeline on 77 million video control events from 66 online courses. Our findings demonstrate that classifiers based on MLLM embeddings reliably predict interaction peaks, generalize to unseen academic fields, and encode interpretable, theory-relevant instructional concepts. Overall, our results show the feasibility of cost-efficient, interpretable pre-screening of educational video design and open new opportunities to empirically examine multimedia learning theory at scale.

教育技术多模态模型可解释性学习分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。