让大模型评分更符合评分标准,通过诊断错误步骤并修正。
EDIT: Evidence-Diagnosed Intervention Training for Rule-Faithful LLM Grading

- 用模型内部信号定位评分中的错误推理环节。
- 在真实多学科数据集上,评分准确率显著超越现有方法。
- 适合需要严格遵循评分标准的教育评测场景。
可靠的评分不仅要求预测准确,还必须基于评分标准和学生作答证据。现有打分分配与干预方法主要针对数学推理等自包含任务,在此场景下因无法识别评分逻辑出错位置或模型信念变化过程而表现不佳。我们提出证据诊断式干预训练(EDIT),一种两阶段框架,用于训练更符合评分标准的大模型评分器。第一阶段,EDIT-SFT 利用模型内部信号——最终分数的后验信念与输入关联得分——定位问题推理步骤,并借助评分清单修正这些局部步骤。第二阶段,EDIT-RL 通过信念引导的奖励塑造校准评分器,在惩罚有害信念偏移的同时保留有益探索空间。在两个真实世界、跨学科评分基准上的实验表明,EDIT 在域内与域外划分上均持续优于强监督微调与强化学习基线,消融实验证实内部状态诊断是性能提升的关键。
原文摘要 · Abstract (English)
Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student answer. Existing credit-assignment and intervention methods, primarily designed for self-contained reasoning tasks such as mathematics reasoning, struggle in this setting because they do not identify where grading reasoning goes wrong or how the model's belief about the final mark changes during reasoning. We propose Evidence-Diagnosed Intervention Training (EDIT), a two-phase framework for training more rubric-faithful LLM graders. First, EDIT-SFT locates problematic reasoning steps using internal model signals: posterior belief over the final mark and input-grounding scores. It then revises only these local steps with help from a rubric checklist. Second, EDIT-RL calibrates the grader with belief-guided reward shaping, penalising large harmful belief drifts while still allowing helpful exploration. Experiments on two real-world, multi-subject grading benchmarks demonstrate that EDIT consistently outperforms strong supervised fine-tuning and reinforcement learning baselines on both in-domain and out-of-domain splits, with ablation studies confirming that internal-state diagnostics drive these gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。