自动发现大模型评估中的潜在偏见,提升评测可靠性。
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
- 用大模型驱动框架自动探测评估中的未知偏见。
- 在JudgeBench-Pro上,强模型错误率超50%。
- 适合关注评测公平性与模型鲁棒性的研究者。
LLM-as-a-Judge已被广泛应用于各类研究与实际场景,但其评估的稳健性和可靠性仍是关键问题。核心挑战在于偏见,现有研究多聚焦已知偏见及其影响,而对潜在未知偏见的自动化、系统性探索仍不足。为此,我们提出BiasScope,一个由大模型驱动的框架,可大规模、自动化发现模型评估中可能存在的潜在偏见。该框架在JudgeBench数据集上验证了其通用性与有效性,能覆盖不同模型家族与规模。它突破了传统方法依赖人工和预设偏见列表的局限,将偏见发现从被动转为主动、全面的自动化过程。基于BiasScope,我们进一步提出JudgeBench-Pro,作为更难的评测基准,用于评估LLM-as-a-judge的鲁棒性。令人震惊的是,即使强大模型在JudgeBench-Pro上的错误率也超过50%,凸显加强评估鲁棒性与缓解潜在偏见的紧迫性。
原文摘要 · Abstract (English)
LLM-as-a-Judge has been widely adopted across various research and practical applications, yet the robustness and reliability of its evaluation remain a critical issue. A core challenge it faces is bias, which has primarily been studied in terms of known biases and their impact on evaluation outcomes, while automated and systematic exploration of potential unknown biases is still lacking. Nevertheless, such exploration is crucial for enhancing the robustness and reliability of evaluations. To bridge this gap, we propose BiasScope, a LLM-driven framework for automatically and at scale discovering potential biases that may arise during model evaluation. BiasScope can uncover potential biases across different model families and scales, with its generality and effectiveness validated on the JudgeBench dataset. It overcomes the limitations of existing approaches, transforming bias discovery from a passive process relying on manual effort and predefined bias lists into an active and comprehensive automated exploration. Moreover, based on BiasScope, we propose JudgeBench-Pro, an extended version of JudgeBench and a more challenging benchmark for evaluating the robustness of LLM-as-a-judge. Strikingly, even powerful LLMs as evaluators show error rates above 50\% on JudgeBench-Pro, underscoring the urgent need to strengthen evaluation robustness and to mitigate potential biases further.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。