测试大模型评判者对不确定表达的敏感度,发现其易受质疑语气影响。
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
- 构建 EMBER 基准,评估大模型在单/双样本评价中对认知标记的鲁棒性。
- 所有测试模型(含 GPT-4o)均表现出对不确定表达的显著负面偏见。
- 模型更排斥表达不确定性的标记,而非仅关注内容正确性。
为遵循诚实原则,训练大语言模型生成包含认知标记的输出正日益普遍。然而,带有认知标记的评价却长期被忽视,引发关键问题:认知标记是否会导致大模型评判结果产生意外负面影响?为此,我们提出 EMBER 基准,用于评估大模型评判者在单样本与双样本评价设置下对认知标记的鲁棒性。基于 EMBER 的评估结果表明,所有测试的大模型评判者(包括 GPT-4o)均表现出明显缺乏鲁棒性。具体而言,模型对认知标记存在负向偏见,尤其针对表达不确定性的标记更为显著,说明大模型评判者会受到标记本身的影响,未能专注于内容正确性。
原文摘要 · Abstract (English)
In line with the principle of honesty, there has been a growing effort to train large language models (LLMs) to generate outputs containing epistemic markers. However, evaluation in the presence of epistemic markers has been largely overlooked, raising a critical question: Could the use of epistemic markers in LLM-generated outputs lead to unintended negative consequences? To address this, we present EMBER, a benchmark designed to assess the robustness of LLM-judges to epistemic markers in both single and pairwise evaluation settings. Our findings, based on evaluations using EMBER, reveal that all tested LLM-judges, including GPT-4o, show a notable lack of robustness in the presence of epistemic markers. Specifically, we observe a negative bias toward epistemic markers, with a stronger bias against markers expressing uncertainty. This suggests that LLM-judges are influenced by the presence of these markers and do not focus solely on the correctness of the content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。