对齐导致大模型评估时数值偏倚,影响评分公正性。
Exploring the Effects of Alignment on Numerical Bias in Large Language Models
- 对比对齐前后模型,发现对齐加剧数值偏倚
- 调整评分范围可有效降低偏倚,提升评估性能
- 适合关注大模型评估公平性的研究人员
将大语言模型(LLM)用作评估者(LLM-as-a-judge)在诸多评测任务中表现优异。然而,这类评估模型存在数值偏倚现象,即某些评分被生成的频率远高于其他分数,导致评估性能下降。本研究探究该偏倚成因。由于多数评估型LLM通过指令微调和偏好对齐进行训练,且已有研究指出对齐会降低输出多样性,我们提出假设:数值偏倚源于对齐过程。通过比较对齐前后的模型输出,结果表明对齐确实加剧了数值偏倚。此外,本文尝试温度调节、分布校准和评分范围调整等缓解策略,其中评分范围调整效果最佳,但仍属启发式方法。研究强调需进一步探索最优评分范围选择及更稳健的缓解机制。
原文摘要 · Abstract (English)
"LLM-as-a-judge," which utilizes large language models (LLMs) as evaluators, has proven effective in many evaluation tasks. However, evaluator LLMs exhibit numerical bias, a phenomenon where certain evaluation scores are generated disproportionately often, leading reduced evaluation performance. This study investigates the cause of this bias. Given that most evaluator LLMs are aligned through instruction tuning and preference tuning, and that prior research suggests alignment reduces output diversity, we hypothesize that numerical bias arises from alignment. To test this, we compare outputs from pre- and post-alignment LLMs, and observe that alignment indeed increases numerical bias. We also explore mitigation strategies for post-alignment LLMs, including temperature scaling, distribution calibration, and score range adjustment. Among these, score range adjustment is most effective in reducing bias and improving performance, though still heuristic. Our findings highlight the need for further work on optimal score range selection and more robust mitigation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。