用大模型评估生成文本质量,发现其易受干扰且不如人工可靠。
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
- 用谷歌Gemini 1做自动评分,要求给出分数和理由。
- 在SummEval和USR数据集上,模型与人类评分一致性有限。
- 输入微调后模型表现明显下降,不适合独立使用。
传统评价指标如BLEU和ROUGE难以捕捉生成文本的细微品质,尤其在缺乏唯一标准答案时。本文探索大型语言模型(LLM)——以Google Gemini 1为例——在摘要和对话任务中作为非标准化评价指标的自动评估器的潜力。我们在SummEval和USR数据集上,通过多种提示策略对比大模型评分与人类判断的一致性,并要求模型同时输出评分及理由。此外,我们测试了大模型评估器在输入扰动下的鲁棒性。结果表明,尽管大模型展现出一定潜力,但其与人类评价者的对齐程度有限,对扰动不鲁棒,距离独立作为主观评价指标的可靠工具仍有显著差距。
原文摘要 · Abstract (English)
Traditional evaluation metrics like BLEU and ROUGE fall short when capturing the nuanced qualities of generated text, particularly when there is no single ground truth. In this paper, we explore the potential of Large Language Models (LLMs), specifically Google Gemini 1, to serve as automatic evaluators for non-standardized metrics in summarization and dialog-based tasks. We conduct experiments across multiple prompting strategies to examine how LLMs fare as quality evaluators when compared with human judgments on the SummEval and USR datasets, asking the model to generate both a score as well as a justification for the score. Furthermore, we explore the robustness of the LLM evaluator by using perturbed inputs. Our findings suggest that while LLMs show promise, their alignment with human evaluators is limited, they are not robust against perturbations and significant improvements are required for their standalone use as reliable evaluators for subjective metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。