教育AI标注别只看人工一致性,更要关注教学实效性。
Beyond Agreement: Rethinking Ground Truth in Educational AI Annotation
- 用多标签、专家评审等方法替代单一共识度评估
- 实证显示新方法能更好预测学生学习效果
- 适合教育数据建模与智能辅导系统研发者
人类评估者常存在偏见、不可靠,难以定义真正‘真实值’。尽管教育AI需大量训练数据,传统互评可靠性(IRR)指标如Cohen's kappa仍被广泛使用。本文指出,过度依赖人工一致性会阻碍有效数据分类,进而影响学习改进。为此,提出五种互补评估方法:多标签标注、专家指导、闭环验证等,强调外部效度——例如在多种辅导行为(如提供提示)中验证辅导动作的有效性。这些方法更有利于生成提升学习成效、带来可操作洞见的模型。呼吁学界重新定义标注质量,从追求共识转向重视有效性与教育影响。
原文摘要 · Abstract (English)
Humans can be notoriously imperfect evaluators. They are often biased, unreliable, and unfit to define "ground truth." Yet, given the surging need to produce large amounts of training data in educational applications using AI, traditional inter-rater reliability (IRR) metrics like Cohen's kappa remain central to validating labeled data. IRR remains a cornerstone of many machine learning pipelines for educational data. Take, for example, the classification of tutors' moves in dialogues or labeling open responses in machine-graded assessments. This position paper argues that overreliance on human IRR as a gatekeeper for annotation quality hampers progress in classifying data in ways that are valid and predictive in relation to improving learning. To address this issue, we highlight five examples of complementary evaluation methods, such as multi-label annotation schemes, expert-based approaches, and close-the-loop validity. We argue that these approaches are in a better position to produce training data and subsequent models that produce improved student learning and more actionable insights than IRR approaches alone. We also emphasize the importance of external validity, for example, by establishing a procedure of validating tutor moves and demonstrating that it works across many categories of tutor actions (e.g., providing hints). We call on the field to rethink annotation quality and ground truth--prioritizing validity and educational impact over consensus alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。