arXiv:2604.21137cs.CLcs.AI2026-04

用联合多任务学习自动分析科学课堂对话中的推理成分,提升教学研究效率。

Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification

论文配图:Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification
图 1 · 摘自论文原文
  • 联合分类教师与学生话语的类型和推理成分,采用分层重采样与大模型生成数据增强。
  • 在少数类识别上,大模型增广使分类性能显著提升,推理成分任务对简单模型也有效。
  • 适合教育技术、认知科学及智能教学系统研究者使用,尤其关注课堂互动分析。

分析科学课堂中学生的推理模式对理解知识建构机制和优化教学实践至关重要,但大规模人工标注对话仍极为耗时。本文提出自动化对话分析系统(ADAS),联合分类教师与学生话语的两个互补维度:话语类型(UT)与基于先前CDAT框架的推理成分(RC)。为缓解少数类别严重不平衡问题,我们采用(1)分层重采样语料库,(2)基于大模型的合成数据增强聚焦少数类,(3)训练双探针头的RoBERTa-base分类器。零样本GPT-5.4基线在UT和RC任务上的宏平均F1分别为0.467和0.476,确立了仅靠提示方法的合理上限,推动后续微调。除分类外,还开展话语模式分析,包括UT×RC共现分析、每会话的认知复杂度指数(CCI)计算、滞后序列分析及IRF链分析,结果显示教师‘带提问的反馈’(Fq)动作是学生推理性推理(SR-I)最稳定的前导因素。结果表明,大模型增强显著提升少数类识别能力,且推理成分任务结构简单,即使词汇基线也能处理。

原文摘要 · Abstract (English)

Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and improving instructional practice to maximize cognitive engagement, yet manual coding of classroom discourse at scale remains prohibitively labor-intensive. We present an automated discourse analysis system (ADAS) that jointly classifies teacher and student utterances along two complementary dimensions: Utterance Type and Reasoning Component derived from our prior CDAT framework. To address severe label imbalance among minority classes, we (1) stratify-resplit the annotated corpus, (2) apply LLM-based synthetic data augmentation targeting minority classes, and (3) train a dual-probe head RoBERTa-base classifier. A zero-shot GPT-5.4 baseline achieves macro-F1 of 0.467 on UT and 0.476 on RC, establishing meaningful upper bounds for prompt-only approaches motivating fine-tuning. Beyond classification, we conduct discourse pattern analyses including UTxRC co-occurrence profiling, Cognitive Complexity Index (CCI) computation per session, lag-sequential analysis, and IRF chain analysis, revealing that teacher Feedback-with-Question (Fq) moves are the most consistent antecedents of student inferential reasoning (SR-I). Our results demonstrate that LLM-based augmentation meaningfully improves UT minority-class recognition, and that the structural simplicity of the RC task makes it tractable even for lexical baselines.

课堂分析推理识别多任务学习教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。