提出可解释评分框架,让AI评分既准又透明。
Interpretability from the Ground Up: Stakeholder-Centric Design of Automated Scoring in Educational Assessments
- 基于四原则设计可解释评分框架,聚焦用户需求。
- 在ASAP-SAS数据集上接近顶尖模型精度,误差仅0.06 QWK。
- 适合教育评估中需透明、可信评分的场景。
基于人工智能的自动化评分系统能高效评估学生生成的复杂回答。然而,尽管对透明性和可解释性的需求日益增长,该领域尚未形成广泛接受的可解释评分解决方案。本文从使用者视角出发,分析不同评估利益相关方的需求与潜在收益,提出面向这些需求的四项可解释性原则:(F)忠实性、(G)全面性、(T)可追溯性、(I)可替换性(FGTI)。为验证其可行性,我们构建了AnalyticScore框架。应用于文本类构造性回答评分时,该框架在评分准确率上优于多种不可解释方法,并在10个ASAP-SAS数据集项目上平均仅比最先进不可解释模型低0.06 QWK。通过与人类标注者执行相同特征提取任务的对比,进一步证明AnalyticScore的特征提取行为与人类高度一致。
原文摘要 · Abstract (English)
AI-driven automated scoring systems offer scalable and efficient means of evaluating complex student-generated responses. Yet, despite increasing demand for transparency and interpretability, the field has yet to develop a widely accepted solution for interpretable automated scoring to be used in large-scale real-world assessments. This work takes a principled approach to address this challenge. We analyze the needs and potential benefits of interpretable automated scoring for various assessment stakeholder groups and develop four principles of interpretability -- (F)aithfulness, (G)roundedness, (T)raceability, and (I)nterchangeability (FGTI) -- targeted at those needs. To illustrate the feasibility of implementing these principles, we develop the AnalyticScore framework as a reference framework. When applied to the domain of text-based constructed-response scoring, AnalyticScore outperforms many uninterpretable scoring methods in terms of scoring accuracy and is, on average, within 0.06 QWK of the uninterpretable SOTA across 10 items from the ASAP-SAS dataset. By comparing against human annotators conducting the same featurization task, we further demonstrate that the featurization behavior of AnalyticScore aligns well with that of humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。