让自动评分模型同时具备排名和打分能力,提升实用性和兼容性。
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
- 用合成摘要模拟人工对比,实现测试时的直接打分
- 在三个评测集上表现接近顶尖对比评估方法
- 适合需要阈值判断的文本生成质量评估场景
随着大语言模型被广泛用于自由文本内容(如摘要、对话、故事生成)的自动评分,研究者致力于衡量其与人类判断的相关性。对于样本级性能评估,基于机器生成文本对之间对比的方法表现良好,但通常无法为单个摘要分配绝对分数,而这在需要阈值判断的应用中至关重要。本文提出一种直接评分方法,利用合成摘要在测试时充当机器对比排序。结果显示,该方法在SummEval(+0.03)、TopicalChat(-0.03)和HANNA(+0.05)三个元评估基准上的轴向平均样本级相关性,与当前最先进对比评估方法相当,并公开了合成上下文摘要数据以支持后续研究。
原文摘要 · Abstract (English)
As large-language models have been increasingly used as automatic raters for evaluating free-form content, including document summarization, dialog, and story generation, work has been dedicated to evaluating such models by measuring their correlations with human judgment. For \textit{sample-level} performance, methods which operate by using pairwise comparisons between machine-generated text perform well but often lack the ability to assign absolute scores to individual summaries, an ability crucial for use cases that require thresholding. In this work, we propose a direct-scoring method which uses synthetic summaries to act as pairwise machine rankings at test time. We show that our method performs comparably to state-of-the-art pairwise evaluators in terms of axis-averaged sample-level correlations on the SummEval (\textbf{+0.03}), TopicalChat (\textbf{-0.03}), and HANNA (\textbf{+0.05}) meta-evaluation benchmarks, and release the synthetic in-context summaries as data to facilitate future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。