用多模态数据提升大模型对中学生协作解题能力的诊断精度。
Rethinking the Potential of Multimodality in Collaborative Problem Solving Diagnosis with Large Language Models
- 融合文本与语音嵌入的多模态模型提升协作解题行为识别效果。
- 在社会认知类解题指标上,多模态比单模态性能提升显著。
- 适用于特定复杂度指标,需结合模型架构与数据特点选择方法。
从数字痕迹中检测协作与问题解决行为以评估学生的协作问题解决(CPS)能力,是人工智能教育(AIEd)领域的长期目标。尽管多模态数据与先进模型被认为有潜力识别复杂的CPS行为,但其实际价值仍缺乏充分实证支持,且存在矛盾结果。本研究在真实教育场景中,针对78名中学生的CPS子技能与指标,探究多模态数据对诊断模型性能的提升潜力。具体采用口语数据的文本嵌入与音频数据的声学嵌入,构建多模态分类模型进行CPS诊断。结果表明,基于Transformer的多模态模型优于传统模型;虽多模态未提升传统单模态模型表现,但在基于Transformer的模型中,其在社会认知类CPS指标上的诊断效果优于单模态模型。研究指出,多模态与建模技术的选择并非万能,其有效性受限于特定类型的CPS指标,受标签复杂度与数据集指标构成影响。最后,论文强调在自动化CPS诊断中应重视人机互补,并建议探索更适配的模型架构与技术,以提升真实教育场景中的诊断能力。
原文摘要 · Abstract (English)
Detecting collaborative and problem-solving behaviours from digital traces to interpret students' collaborative problem solving (CPS) competency is a long-term goal in the Artificial Intelligence in Education (AIEd) field. Although multimodal data and advanced models are argued to have the potential to detect complex CPS behaviours, empirical evidence on their value remains limited with some contrasting evidence. In this study, we investigated the potential of multimodal data to improve model performance in diagnosing 78 secondary school students' CPS subskills and indicators in authentic educational settings. In particular, text embeddings from verbal data and acoustic embeddings from audio data were used in a multimodal classification model for CPS diagnosis. Both unimodal and multimodal transformer-based models outperformed traditional models in detecting CPS classes. Although the inclusion of multimodality did not improve the performance of traditional unimodal models, its integration into transformer-based models demonstrated improved performance for diagnosing social-cognitive CPS classes compared to unimodal transformer-based models. Based on the results, the paper argues that multimodality and the selection of a particular modelling technique should not be taken for granted to achieve the best performance in the automated detection of every CPS subskill and indicator. Rather, their value is limited to certain types of CPS indicators, affected by the complexity of the labels, and dependent on the composition of indicators in the dataset. We conclude the paper by discussing the required nuance when considering the value of LLMs and multimodality in automated CPS diagnosis, highlighting the need for human-AI complementarity, and proposing the exploration of relevant model architectures and techniques to improve CPS diagnosis in authentic educational contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。