arXiv:2605.27402cs.CYcs.AI2026-05

让自动评分更透明可信,可解释且符合评分标准。

REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading

论文配图:REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
图 1 · 摘自论文原文
  • 用可解释的概念模型融合评分标准,提升可读性。
  • 在多个数据集上准确率优于现有方法,概念推理更真实。
  • 适合教育工作者审核与干预自动评分结果。

开放式评分对实现公平与个性化教育至关重要,但人工评分耗时费力,亟需自动化系统。尽管基于神经网络和大语言模型的系统表现优异,但其评分过程如同黑箱,难以验证与信任。概念瓶颈模型(CBMs)通过引入人类可理解的概念,提供机制层面的透明保障。然而,传统CBMs未针对开放式评分优化:未能显式建模细粒度评分标准,忽略评分等级的序数语义,且忽视人工标注概念中的固有可靠性问题。为此,我们提出REC-CBM——一种面向评分标准的误差纠正概念瓶颈模型。该模型引入评分标准感知的概念编码器,学习回答中概念的特定表示,并采用序数成对校准目标,保留评分维度间的排序结构。同时,设计潜在概念误差纠正模块,在最终评分前净化概念预测,同时保持可解释性。在公开数据集上的实验表明,REC-CBM持续提升评分性能,生成更忠实的概念级推理,优于当前最佳基线。进一步分析验证各组件贡献,并证明其在真实教育场景中的适用性。整体而言,本工作提供了一种实用、可解释的评分方案,使教育者能审查、干预并信任自动化决策,推动更透明、可信的教育发展。

原文摘要 · Abstract (English)

Open-ended grading is central to equitable and personalized education, yet manual grading remains time-consuming and costly, underscoring the need for automated grading systems. Although recent neural and large language model (LLM) based systems have demonstrated superior performance, they are typically black-box models whose scoring processes and rationales are difficult for educators to verify and trust. Concept bottleneck models (CBMs) have emerged as a promising approach by routing predictions through human-interpretable concepts, providing a mechanistic guarantee of transparency. However, standard CBMs are not tailored to open-ended grading: they do not explicitly model fine-grained rubric dimensions, inadequately capture the ordinal semantics of scoring scales, and neglect inherent reliability issues in human concept annotations. To address these limitations, we propose REC-CBM, a rubric-aware error-correction concept bottleneck model for trustworthy open-ended grading. REC-CBM introduces a rubric-aware concept encoder that learns concept-specific representations over responses and an ordinal pairwise calibration objective that preserves ranking structure among rubric dimensions. It further incorporates a latent concept error-correction module that denoises concept predictions before final grade prediction while preserving interpretability. Comprehensive experiments on publicly available datasets show that REC-CBM consistently improves grading performance and produces more faithful concept-level reasoning than both state-of-the-art baselines. Further analyses validate the contribution of each component and demonstrate the applicability in realistic educational settings. Overall, this work provides a practical, interpretable grading solution that enables educators to inspect, intervene in, and trust automated decisions, advancing more transparent and trustworthy education.

自动评分可解释性教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。