发现大模型评判时偏好熟悉文本,可能误导评估结果
Self-Preference Bias in LLM-as-a-Judge
- 提出用困惑度衡量大模型自偏见,量化评判偏差
- GPT-4对低困惑度输出评分显著更高,无论是否自生成
- 适合关注AI评估可信性、模型偏见的研究者
基于大语言模型(LLM)的自动化评估在对话系统性能评测中广泛应用,但其自偏见问题带来显著风险,可能导致特定风格或策略被过度推崇。现有方法缺乏对自偏见的定量测量,且成因不明。本文提出一种新型量化指标,实验表明GPT-4存在显著自偏见。我们假设大模型偏好更熟悉的输出(以更低困惑度为标志),分析发现:无论输出是否由模型自动生成,大模型对低困惑度文本的评分显著高于人类评价者,说明自偏见本质源于对熟悉文本的偏好。
原文摘要 · Abstract (English)
Automated evaluation leveraging large language models (LLMs), commonly referred to as LLM evaluators or LLM-as-a-judge, has been widely used in measuring the performance of dialogue systems. However, the self-preference bias in LLMs has posed significant risks, including promoting specific styles or policies intrinsic to the LLMs. Despite the importance of this issue, there is a lack of established methods to measure the self-preference bias quantitatively, and its underlying causes are poorly understood. In this paper, we introduce a novel quantitative metric to measure the self-preference bias. Our experimental results demonstrate that GPT-4 exhibits a significant degree of self-preference bias. To explore the causes, we hypothesize that LLMs may favor outputs that are more familiar to them, as indicated by lower perplexity. We analyze the relationship between LLM evaluations and the perplexities of outputs. Our findings reveal that LLMs assign significantly higher evaluations to outputs with lower perplexity than human evaluators, regardless of whether the outputs were self-generated. This suggests that the essence of the bias lies in perplexity and that the self-preference bias exists because LLMs prefer texts more familiar to them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。