arXiv:2505.07902cs.CYcs.AI2025-05被引 10

用多模态注意力模型自动评估课堂对话质量,助力教师教学改进。

Multimodal Assessment of Classroom Discourse Quality: A Text-Centered Attention-Based Multi-Task Learning Approach

  • 基于文本中心的多模态融合与多任务学习,联合评估三类课堂对话质量。
  • 在92节德国数学课数据上,综合评分达0.384(接近人工一致性0.326)。
  • 文本信息主导模型表现,音频特征提升评分一致性,适合教育AI研究者。

课堂对话是教与学的核心载体。评估对话实践的不同特征并将其与学生学业成就关联,有助于理解教学质量。传统评估依赖人工编码观察协议,耗时且成本高。尽管已有研究利用AI技术分析单个发言,但对整节课段对话实践的评价仍有限。为此,本研究提出一种以文本为中心的多模态融合架构,基于全球教学洞察(GTI)观察协议,评估三类对话质量:对话性质、提问和解释。首先,采用注意力机制捕捉转录文本、音频和视频流之间的跨模态与内部交互;其次,使用多任务学习联合预测三类成分的质量得分;第三,将任务定义为序数分类问题以体现评分等级顺序。在包含92节录像数学课的GTI Germany数据集上,消融实验验证了设计的有效性。结果表明,文本模态在该任务中起主导作用。整合声学特征可提升模型与人工评分的一致性,整体获得0.384的加权二次κ值,接近人工评分者间可靠性(0.326)。本研究为未来自动化对话质量评估奠定了基础,支持通过及时反馈促进教师专业发展。

原文摘要 · Abstract (English)

Classroom discourse is an essential vehicle through which teaching and learning take place. Assessing different characteristics of discursive practices and linking them to student learning achievement enhances the understanding of teaching quality. Traditional assessments rely on manual coding of classroom observation protocols, which is time-consuming and costly. Despite many studies utilizing AI techniques to analyze classroom discourse at the utterance level, investigations into the evaluation of discursive practices throughout an entire lesson segment remain limited. To address this gap, our study proposes a novel text-centered multimodal fusion architecture to assess the quality of three discourse components grounded in the Global Teaching InSights (GTI) observation protocol: Nature of Discourse, Questioning, and Explanations. First, we employ attention mechanisms to capture inter- and intra-modal interactions from transcript, audio, and video streams. Second, a multi-task learning approach is adopted to jointly predict the quality scores of the three components. Third, we formulate the task as an ordinal classification problem to account for rating level order. The effectiveness of these designed elements is demonstrated through an ablation study on the GTI Germany dataset containing 92 videotaped math lessons. Our results highlight the dominant role of text modality in approaching this task. Integrating acoustic features enhances the model's consistency with human ratings, achieving an overall Quadratic Weighted Kappa score of 0.384, comparable to human inter-rater reliability (0.326). Our study lays the groundwork for the future development of automated discourse quality assessment to support teacher professional development through timely feedback on multidimensional discourse practices.

课堂评估多模态学习教育AI注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。