发现大模型评分会受偏见影响,提出四类缓解方法。
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
- 对比6种大模型评委,识别11类显隐性评价偏见。
- 用有偏数据微调会显著降低模型评分准确性。
- 任务越难评分越低,开放推理题得分更高。
大型语言模型(LLMs)正被用于通信系统中自动评估内容质量,如电信客服对话的响应评价。然而,这些AI“评委”的公正性无法保证,其评价标准中的偏见可能扭曲结果并损害用户信任。本文系统研究了6种基于提示和微调的大模型评委在点对点评分设置下的判断偏见,涵盖11类偏见,包括显性和隐性形式。我们发现,当前最先进的模型对有偏输入具有鲁棒性,通常给予比对应无偏样本更低的分数。进一步研究表明,使用高分但有偏的响应进行微调会显著降低模型性能,凸显训练数据偏见的风险。此外,我们观察到评分与任务难度相关:如GPQA这类难题集平均得分较低,而开放推理数据集(如JudgeLM-val)则平均得分较高。最后,我们提出了四种潜在的缓解策略,以确保实际通信场景中人工智能评判的公平性与可靠性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly being used to autonomously evaluate the quality of content in communication systems, e.g., to assess responses in telecom customer support chatbots. However, the impartiality of these AI "judges" is not guaranteed, and any biases in their evaluation criteria could skew outcomes and undermine user trust. In this paper, we systematically investigate judgment biases across 6 LLM-as-a-judge models spanning both prompt-based and fine-tuned judges under the pointwise scoring setting, encompassing 11 types of biases that cover both implicit and explicit forms. We observed that state-of-the-art LLM judges demonstrate robustness to biased inputs, generally assigning them lower scores than the corresponding clean samples. We further found that fine-tuning an LLM on high-scoring yet biased responses can significantly degrade its performance, highlighting the risk of training on biased data. We also discovered that the judged scores correlate with task difficulty: a challenging dataset like GPQA yields lower average scores, whereas an open-ended reasoning dataset (e.g., JudgeLM-val) sees higher average scores. Finally, we proposed four potential mitigation strategies to ensure fair and reliable AI judging in practical communication scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。