让大模型自动生成题目、作答并评分时,如何确保评分可信?
Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment

- 用生成与评分的一致性衡量评分可靠性,发现模型仅恢复一半预期能力差异
- 语法类技能评分准确度高(相关系数0.7以上),设计类技能几乎无效
- 低水平错误被高估,可能误导学生分组,需细化评分标准来改善
当同一大型语言模型同时生成测评题、模拟学生作答并评分时,验证过程陷入自我循环。本文提出生成-评价一致性(Generative-Evaluative Agreement, GEA)作为有效性标准,衡量模型评分是否能还原其生成时设定的能力水平。在一项两阶段自适应评估中首次直接测量GEA,结果显示模型仅恢复约一半预期方差(相关系数r=0.698),且存在系统性正向偏差。对于可语法验证的技能,GEA较强(r>0.7);而对设计层级技能,GEA接近零。低能力水平的错误被过度高估,导致分数在分流阈值附近虚高。本文主张采用细粒度、技能分解的评分标准作为主要改进机制,并提出补充缓解策略。
原文摘要 · Abstract (English)
When the same LLM generates assessment items, simulates student responses, and scores them, the validation loop is self-referential. We introduce Generative-Evaluative Agreement (GEA), a validity criterion measuring whether an LLM's scoring function recovers the skill levels its generative function was instructed to produce. In the first direct measurement of GEA on a two-stage adaptive assessment, the model recovers roughly half the intended variance r = 0.698 with systematic positive bias. GEA is strong r > 0.7 for syntactically verifiable skills but near zero for design-level skills, and low-skill overestimation inflates scores near the routing threshold. We argue that granular, skill-decomposed rubrics are the principal proposed mechanism for strengthening GEA and outline complementary mitigations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。