研究导师如何评分,发现AI评分权重需调整才能更贴近人工判断。
AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment

- 通过调查84位导师,获取35项评分标准的实际权重。
- 调整AI权重后,评分偏差从11.18%降至10.85%,但提升不显著。
- 导师间评分一致性远高于AI与导师的匹配度,适合教育评估优化者参考。
基于评分量表的AI论文评估系统通过权重分配不同评价标准的重要性,这些权重通常由专家设定,但缺乏关于导师实际优先级的实证依据。本研究调查了四个学科领域共84位论文导师,收集了35项论文评估标准的权重数据。与AI评估系统RubiSCoT的默认权重对比发现,导师权重存在显著差异。为评估实际影响,将导师权重整合至多个校准配置,并在80篇德语论文上测试。最佳配置将AI生成评分与导师评分间的平均相对偏差从11.18%降至10.85%,但改善未达统计显著。人类导师间的平均相互偏差仅为4.44%。结果表明,仅靠准则权重校准无法显著提升AI与人工评估的一致性。
原文摘要 · Abstract (English)
Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although little empirical evidence exists regarding how thesis supervisors actually prioritize evaluation criteria. Consequently, this study investigates supervisor-derived criterion weights in thesis assessment and evaluates their impact on AI-based assessment. We surveyed 84 thesis supervisors across four academic disciplines and collected weighting data for 35 thesis assessment criteria. Comparison with the default criterion weights of the AI assessment system RubiSCoT [1] revealed substantial divergences between supervisor-derived and default criterion weights. To evaluate the practical implications of these differences, the supervisor-derived weights were integrated into multiple calibration configurations and evaluated on a corpus of 80 German-language theses. The best-performing configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, although the improvement was not statistically significant. Human supervisors showed substantially stronger agreement with each other, exhibiting a mean inter-supervisor relative deviation of 4.44%. The findings indicate that criterion-weight calibration alone does not substantially improve alignment between AI-generated and human assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。