arXiv:2507.14579cs.CLcs.AI2025-07中稿 · appear in the work…

用多模态BERT提升教育AI对协作问题解决的诊断,发现人机互补关键在模型可解释性。

Exploring Human-AI Complementarity in CPS Diagnosis Using Unimodal and Multimodal BERT Models

  • 融合语音与声学特征的AudiBERT模型优于传统BERT,尤其在社会认知维度上
  • 数据量越大,召回率越高;人类标注一致性影响BERT精确率
  • 模型可解释性是实现人机协同诊断的核心,助力教师反思性编码

利用机器学习检测协作问题解决(CPS)指标是教育人工智能的重要挑战。近期研究使用双向编码器表示变换器(BERT)模型处理转录文本,有效识别有意义的CPS指标。一项重要进展是多模态BERT变体AudiBERT,它整合了语音与声学-韵律特征,提升了CPS诊断效果。尽管初步结果表明多模态有优势,但统计显著性不明确,且缺乏对人机互补策略的指导。本研究发现,AudiBERT不仅改善了数据稀疏类别的分类表现,还在社会认知维度上显著优于BERT模型;但在情感维度未见类似提升。相关分析显示,训练数据量与AudiBERT和BERT的召回率显著正相关;而BERT的精确率与人工标注者间高一致性显著相关。当使用BERT诊断AudiBERT已准确识别的子技能时,整体表现不一致。论文最后提出结构化的人机互补框架,强调模型可解释性对于支持教师在反思性编码过程中的自主性和参与度至关重要。

原文摘要 · Abstract (English)

Detecting collaborative problem solving (CPS) indicators from dialogue using machine learning techniques is a significant challenge for the field of AI in Education. Recent studies have explored the use of Bidirectional Encoder Representations from Transformers (BERT) models on transcription data to reliably detect meaningful CPS indicators. A notable advancement involved the multimodal BERT variant, AudiBERT, which integrates speech and acoustic-prosodic audio features to enhance CPS diagnosis. Although initial results demonstrated multimodal improvements, the statistical significance of these enhancements remained unclear, and there was insufficient guidance on leveraging human-AI complementarity for CPS diagnosis tasks. This workshop paper extends the previous research by highlighting that the AudiBERT model not only improved the classification of classes that were sparse in the dataset, but it also had statistically significant class-wise improvements over the BERT model for classifications in the social-cognitive dimension. However, similar significant class-wise improvements over the BERT model were not observed for classifications in the affective dimension. A correlation analysis highlighted that larger training data was significantly associated with higher recall performance for both the AudiBERT and BERT models. Additionally, the precision of the BERT model was significantly associated with high inter-rater agreement among human coders. When employing the BERT model to diagnose indicators within these subskills that were well-detected by the AudiBERT model, the performance across all indicators was inconsistent. We conclude the paper by outlining a structured approach towards achieving human-AI complementarity for CPS diagnosis, highlighting the crucial inclusion of model explainability to support human agency and engagement in the reflective coding process.

人机协同多模态学习教育AI模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。