研究评分标准修改如何影响人工与AI评分的一致性
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
- 通过调整评分标准的呈现方式,提升人与AI评分一致性
- 提供示例和上下文可提高一致率,复杂度越高越易导致分歧
- 适合内容评估、自动评分系统设计者参考
自动评分器(LLM-as-judges)在内容评估和自动化审核中日益普及。然而,针对评分标准修改对人工与自动评分一致性影响的统计分析仍有限。以整体性评价(如‘作文质量’)为主的评分标准因标准模糊或主观性强,易引发理解不一;而分解性评价(如将‘质量’拆分为‘流畅性’和‘结构’)虽可提升个体评分准确性,却可能导致人工与自动评分不一致。研究发现,增加代表性示例、补充上下文信息并减少评分位置偏差,能提升人机评分一致性;而评分标准过于复杂或采用保守聚合方法则会降低一致性。在自动作文评分和指令遵循评估领域,结果表明需根据具体场景和评分标准分析性能表现,以实现更高的人机评分一致率。
原文摘要 · Abstract (English)
Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifications in a rubric presented to both humans and autoraters affect their score agreement. Rubrics that ask for an overall or \emph{holistic} judgment - for example, rating the ``quality'' of an essay - may be inconsistently interpreted due to the complexity or subjectivity of the criteria. Conversely, rubrics can ask for \emph{analytic} judgments, which decompose assessment criteria - for example, ``quality'' into ``fluency'' and ``organization''. While these rubrics can be edited to improve the individual accuracy of both human and automated scoring, this approach may result in disagreement between the two scores, or with the associated holistic judgment. Designing and deploying reliable autoraters requires understanding not just the relationship between human and autorater annotations but how that relationship changes as holistic or analytic judgments are elicited. The results indicate that rubric edits providing representative examples and additional context, and reducing positional bias in the rubric increased human-autorater agreement, while higher rubric complexity and conservative aggregation methods tended to decrease it. The findings from the automatic essay scoring and instruction-following evaluation domains suggest that practitioners should carefully analyze domain- and rubric-specific performance to move towards higher human-autorater agreement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。