量化大模型评判中的12类偏见,揭示其可靠性问题
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
- 设计自动化框架CALM,系统检测大模型评判的12种偏见
- 实验发现先进模型在特定任务中仍存在显著偏见
- 适合关注AI评估可靠性的研究者与开发者参考
LLM-as-a-Judge 广泛用于各类基准评估和模型训练中的监督奖励。然而,尽管在多个领域表现优异,潜在问题仍未充分探讨,削弱了其可靠性与应用范围。为此,我们识别出12种关键潜在偏见,提出一种新的自动化偏见量化框架CALM,通过自动化且基于原则的修改方式,系统量化并分析每类偏见。实验覆盖多个主流语言模型,结果表明:虽然先进模型整体表现良好,但在某些特定任务中仍存在显著偏见。实证结果显示,LLM-as-a-Judge的可靠性仍有提升空间。此外,我们还讨论了这些偏见的显性和隐性影响,并提出可靠应用建议。本工作强调利益相关方需重视这些问题,提醒用户在使用LLM-as-a-Judge时保持谨慎。
原文摘要 · Abstract (English)
LLM-as-a-Judge has been widely utilized as an evaluation method in various benchmarks and served as supervised rewards in model training. However, despite their excellence in many domains, potential issues are under-explored, undermining their reliability and the scope of their utility. Therefore, we identify 12 key potential biases and propose a new automated bias quantification framework-CALM-which systematically quantifies and analyzes each type of bias in LLM-as-a-Judge by using automated and principle-guided modification. Our experiments cover multiple popular language models, and the results indicate that while advanced models have achieved commendable overall performance, significant biases persist in certain specific tasks. Empirical results suggest that there remains room for improvement in the reliability of LLM-as-a-Judge. Moreover, we also discuss the explicit and implicit influence of these biases and give some suggestions for the reliable application of LLM-as-a-Judge. Our work highlights the need for stakeholders to address these issues and remind users to exercise caution in LLM-as-a-Judge applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。