arXiv:2605.16806cs.LGcs.AI2026-05

通过跨模态对齐提升游戏化学习中协作满意度预测精度

Cross-modal Affinity-aligned Multimodal Learning Analytics for Predicting Student Collaboration Satisfaction in Game-Based Learning

论文配图:Cross-modal Affinity-aligned Multimodal Learning Analytics for Predicting Student Collaboration Satisfaction in Game-Based Learning
图 1 · 摘自论文原文
  • 用亲和矩阵显式建模多模态关系,自适应抑制低信息模态
  • 在50名中学生上测试,比单一模态和现有方法更稳定有效
  • 适合做教育智能分析、多模态融合或可解释性研究者

协作式游戏化学习环境为小组知识建构提供了丰富机会,但自动预测学生协作满意度仍具挑战。主要障碍是模态退化:在教育应用中,眼动等单个模态在不同学生群体间信息量不一致,导致基于隐式注意力的融合产生脆弱的多模态表征。本文提出亲和对齐多模态学习分析框架(AAMLA),核心为跨模态亲和引导的模态对齐(CAMA)模块,通过亲和矩阵显式建模模态间关系,并利用对比学习强化跨模态一致性,实现对低信息模态的自适应抑制而不剔除。AAMLA还引入模态特定投影层,将面部动作单元、头部姿态、眼动及交互日志等异构特征映射至统一语义空间后再对齐。在EcoJourneys协作学习环境中对50名中学生的实验表明,该方法在标准与模态退化条件下均显著优于单模态基线和先前交叉注意力方法;SHAP与t-SNE分析证实,CAMA生成了鲁棒且可解释的跨模态表征,适用于学生协作建模。

原文摘要 · Abstract (English)

Collaborative game-based learning environments offer rich opportunities for small-group knowledge construction, yet automatically predicting student collaboration satisfaction remains challenging. A critical barrier is modality degradation: in educational deployments, individual modalities such as eye gaze exhibit inconsistent informativeness across student cohorts, causing implicit attention-based fusion to produce brittle multimodal representations. We propose the Affinity-Aligned Multimodal Learning Analytics (AAMLA) framework, whose core contribution is the Cross-modal Affinity-guided Modality Alignment (CAMA) module, which explicitly models inter-modal relationships via affinity matrices and enforces cross-modal consistency through contrastive learning, enabling adaptive suppression of uninformative modalities without discarding them. AAMLA further applies modality-specific projection layers to map heterogeneous features, including facial action units, head pose, eye gaze, and interaction trace logs, into a unified semantic space prior to alignment. Experiments on 50 middle school students in the EcoJourneys collaborative learning environment demonstrate consistent improvements over unimodal baselines and prior cross-attention approaches under standard and modality degradation conditions, with SHAP and t-SNE analyses confirming that CAMA produces robust, interpretable cross-modal representations for student collaboration modeling.

多模态学习教育分析可解释性游戏化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。