用淘汰赛式迭代对比,让大模型更准地评分。
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
- 通过多轮配对比较构建淘汰赛,让模型建立全局评分视角。
- 在大学考试和机器翻译评估中,相关性平均提升0.07。
- 适合需要高精度自动评分的教育与语言任务场景。
大型语言模型(LLMs)在机器翻译、科学领域等已展现良好评估能力。现有基于LLM的评分方法多依赖单次独立评估或一轮配对比较,难以形成全局排名视角。为此,我们提出淘汰赛评估(Knockout Assessment),采用迭代式配对比较的淘汰赛机制,使评判模型逐步建立更全面的评分认知。在两个数据集上对三种LLM的实验表明,该方法显著提升了评分准确性:在大学水平考试评分与机器翻译评估中,与专家评分的皮尔逊相关性平均提升0.07,使模型评分更贴近人类判断。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown to be effective evaluators across various domains such as machine translations or the scientific domain. Current LLM-as-a-Judge approaches rely mostly on individual assessments or a single round of pairwise assessments, preventing the judge LLM from developing a global ranking perspective. To address this, we present Knockout Assessment, an LLM-asa Judge method using a knockout tournament system with iterative pairwise comparisons. Experiments across three LLMs on two datasets show that knockout assessment improves scoring accuracy, increasing Pearson correlation with expert evaluations by 0.07 on average for university-level exam scoring and machine translation evaluations, aligning LLM assessments more closely with human scoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。