分离学生回答内容与教师评分偏见,提升自动评分透明度。
Disentangling Learning from Judgment: Representation Learning for Open Response Analytics
- 用教师历史作为动态先验,分离评分倾向与内容质量
- 结合教师先验与文本嵌入,AUC达0.815,显著优于纯内容模型
- 揭示评分差异背后的学习证据,助力教学反思
开放式作答在学习中至关重要,但自动评分常混淆学生内容与教师评分习惯。本文提出以分析为核心的框架,将内容信号与评分者倾向分离,使判断过程可观察、可审计。基于去标识化的ASSISTments数学作答数据,我们建模教师历史为动态先验,使用句子嵌入表示文本。通过质心归一化与响应-题目嵌入差异分析,并引入先验显式建模教师效应,降低题目与教师相关混杂因素影响。时间验证的线性模型量化各信号贡献,模型分歧则暴露可定性检视的异常。结果表明,教师先验对评分预测有显著影响;当先验与内容嵌入结合时表现最优(AUC~0.815),仅依赖内容的模型仍优于随机(AUC~0.626)。校正评分者效应后,内容表征特征选择更精准,保留更多语义信息维度,识别出支持理解而非表面表述差异的真实学习证据。该方法提供实用分析流程,将嵌入从普通特征转化为可反思的学习分析工具,帮助教师与研究者审视评分实践与学生思维证据的一致性。
原文摘要 · Abstract (English)
Open-ended responses are central to learning, yet automated scoring often conflates what students wrote with how teachers grade. We present an analytics-first framework that separates content signals from rater tendencies, making judgments visible and auditable via analytics. Using de-identified ASSISTments mathematics responses, we model teacher histories as dynamic priors and represent text with sentence embeddings. We apply centroid normalization and response-problem embedding differences, and explicitly model teacher effects with priors to reduce problem- and teacher-related confounds. Temporally-validated linear models quantify the contributions of each signal, and model disagreements surface observations for qualitative inspection. Results show that teacher priors heavily influence grade predictions; the strongest results arise when priors are combined with content embeddings (AUC~0.815), while content-only models remain above chance but substantially weaker (AUC~0.626). Adjusting for rater effects sharpens the selection of features derived from content representations, retaining more informative embedding dimensions and revealing cases where semantic evidence supports understanding as opposed to surface-level differences in how students respond. The contribution presents a practical pipeline that transforms embeddings from mere features into learning analytics for reflection, enabling teachers and researchers to examine where grading practices align (or conflict) with evidence of student reasoning and learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。