融合屏幕视频与操作序列,提升机器学习游戏中的解题策略识别准确率。
Multimodal Late Fusion Model for Problem-Solving Strategy Classification in a Machine Learning Game
- 用视觉数据和操作序列做多模态晚期融合
- 准确率比单一模态模型提高15%以上
- 适合需要精细评估学习策略的教育场景
机器学习模型广泛用于数字学习环境中的隐蔽评估。现有方法通常依赖抽象的游戏日志数据,可能忽略与学习者认知策略相关的细微行为线索。本文提出一种多模态晚期融合模型,整合基于屏幕录像的视觉数据与结构化的游戏内操作序列,以分类学生的问题解决策略。在一项针对149名中学生参与多点触控教育游戏的试点研究中,该融合模型优于单模态基线模型,分类准确率提升超过15%。结果表明,多模态机器学习在策略敏感型评估与交互式学习环境中的自适应支持方面具有潜力。
原文摘要 · Abstract (English)
Machine learning models are widely used to support stealth assessment in digital learning environments. Existing approaches typically rely on abstracted gameplay log data, which may overlook subtle behavioral cues linked to learners' cognitive strategies. This paper proposes a multimodal late fusion model that integrates screencast-based visual data and structured in-game action sequences to classify students' problem-solving strategies. In a pilot study with secondary school students (N=149) playing a multitouch educational game, the fusion model outperformed unimodal baseline models, increasing classification accuracy by over 15%. Results highlight the potential of multimodal ML for strategy-sensitive assessment and adaptive support in interactive learning contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。