多模型重复评分验证了大模型在简答评分中的可靠性和诊断价值。
Reliability, validity, and diagnostic evidence for multi-model LLM short-answer scoring
- 采用多模型多次评分,基于五个维度评估答案质量。
- 模型间信度高(ICC达0.97以上),且优于传统方法的AUC与F1。
- 适合教育评估中需证据支持的AI评分场景,非完全替代人工。
大型语言模型(LLMs)在教育评分中的应用日益广泛,但单模型单次评估缺乏充分证据。本研究通过重复多模型OCG-PRES引导评分,对短答案评估进行可靠性、有效性及诊断价值分析。使用996条SciEntsBank答题数据,GPT、DeepSeek和Qianwen分别在三次独立运行中,按概念覆盖、关系准确性、推理完整性、矛盾控制和领域相关性五项维度评分。评分结果与官方二分类及五分类标签对比,并与基于回答长度、Jaccard重叠、TF-IDF余弦相似度及传统逻辑回归模型的非LLM基线比较。所有模型重复评分信度均高,ICC(3,k)分别为:GPT 0.977,DeepSeek 0.992,Qianwen 0.981,其中DeepSeek最稳定。GPT在官方标签匹配上表现最优(AUC=0.909),Qianwen更严格,阈值固定为3.0时精度高但召回率低。OCG-PRES评分呈现预期诊断模式,且在AUC与F1上均优于所有非LLM基线。结果表明,重复多模型评分可为大模型辅助短答评分提供可靠的信效度与诊断证据,建议谨慎用于评分辅助,而非替代人工判断。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used or proposed for educational scoring, but single-model and single-run evaluations provide limited evidence for assessment use. Short-answer scoring requires evidence about reliability, validity, severity, diagnostic value, and failure cases. This study evaluated repeated multi-model OCG-PRES guided LLM scoring for short-answer assessment. The analysis used 996 SciEntsBank responses. GPT, DeepSeek, and Qianwen each scored every response across three independent runs using five OCG-PRES dimensions: concept coverage, relation accuracy, reasoning completeness, contradiction control, and domain relevance. Scores were evaluated against official binary and five-category labels and compared with non-LLM baselines based on answer length, Jaccard keyword overlap, TF-IDF cosine similarity, and a combined traditional logistic model. Repeated-run reliability was high for all models, with ICC(3,k) = .977 for GPT, .992 for DeepSeek, and .981 for Qianwen. DeepSeek was the most stable across runs. GPT showed the strongest official-label alignment by AUC (.909), while Qianwen was stricter, with higher precision but lower recall under the fixed threshold = 3.0 rule. OCG-PRES scores followed expected diagnostic patterns across five official categories and outperformed all non-LLM baselines in AUC and F1. Repeated multi-model OCG-PRES scoring provides reliability, validity, and diagnostic evidence for LLM-assisted short-answer scoring. The findings support cautious, evidence-based use as a scoring support tool rather than a replacement for human judgement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。