arXiv:2606.05180cs.CLcs.AI2026-06ACL

用SHAP和LLM结合解释评分模型为何打分,提升教育评估透明度。

From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment

  • 融合SHAP值与LLM生成理由,实现评分细节可解释。
  • SHAP比LLM理由更准确、影响更大且跨模型通用。
  • 适合关注教育评分可信度的研究者与AI评估系统设计者。

自动化评分模型在复杂语言表现(如课堂实录)的评分中日益普及,但通常无法说明评分依据。本文提出一种面向评分任务的句子级可解释性框架,结合模型无关的Shapley值归因与大语言模型生成的理由。基于NCTE语料库中CLASS框架的反馈质量维度进行实例化,该框架可系统比较微调预训练语言模型(PLMs)与提示式大语言模型(LLMs)在评分性能与解释忠实性上的表现。在6000个标注片段上,微调模型在预测准确率上优于提示模型,但存在向中间分数压缩的现象。删除测试显示,SHAP能识别出对预测有可靠影响的句子,其引发的预测变化幅度更大、逻辑更连贯,优于LLM生成的理由。跨模型分析进一步表明,SHAP归因在不同架构间具有强泛化能力,而LLM理由影响有限且不一致。结果表明,SHAP为评分任务提供更忠实、可迁移的解释,所提框架为高风险教育评估及其它基于量规的语言评估任务中的评分模型及其解释提供了严谨评价基础。

原文摘要 · Abstract (English)

Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts, yet they typically provide little insight into why a particular score is produced. We propose a general framework for sentence-level interpretability of rubric-based scoring that combines model-agnostic Shapley-value attributions with rationales generated by large language models (LLMs). Instantiated on the Quality of Feedback dimension of the CLASS framework using the NCTE corpus, the framework enables systematic comparison of fine-tuned pretrained language models (PLMs) and prompted LLMs on both scoring performance and explanation faithfulness. Across 6k annotated transcript segments, fine-tuned PLMs outperform LLMs in prediction accuracy but exhibit label compression toward mid-scale scores. Deletion-based tests show that SHAP identifies sentences that reliably drive model predictions, producing typically larger and more coherent prediction shifts than LLM-generated rationales. Cross-model analyses further reveal that SHAP attributions transfer robustly across architectures, whereas LLM rationales exert limited and inconsistent influence. Overall, the findings demonstrate that SHAP provides more faithful and transferable explanations for rubric-based scoring, and that the proposed framework offers a principled basis for evaluating both scoring models and their explanations in high-stakes educational settings and other rubric-based language assessment tasks.

可解释AI教育评估SHAPLLM理由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。