arXiv:2505.19477cs.AI2025-05EMNLP被引 15

多智能体评判中,辩论机制会放大偏见,而元评判更抗偏见。

Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge

  • 对比辩论与元评判两种多智能体评判框架的偏见表现。
  • 辩论框架在后续轮次中偏见显著增强,元评判则更具抵抗力。
  • 引入去偏代理可有效缓解辩论中的偏见,但对元评判作用有限。

LLM-as-Judge作为人类评估的可扩展替代方案,已广泛用于训练中的奖励信号生成。尽管近期研究探索了多智能体辩论和元评判等扩展方法以提升评估质量,但内在偏见在这些场景中的表现仍不明确。本研究系统分析了四种偏见类型:立场偏见、冗长偏见、思维链偏见和从众偏见。我们在两种主流多智能体LLM-as-Judge框架——Multi-Agent-Debate和LLM-as-Meta-Judge——上进行评估。结果表明,辩论框架在初始辩论后偏见急剧放大,并在后续轮次持续存在;而元评判框架表现出更强的抗偏见能力。进一步考察将PINE(一种领先的单智能体去偏方法)作为无偏代理引入系统的效果,发现其能有效降低辩论场景中的偏见,但在元评判场景中效果较弱。本研究全面揭示了多智能体评判系统中的偏见行为,强调在协作评估中需采取针对性的去偏策略。

原文摘要 · Abstract (English)

LLM-as-Judge has emerged as a scalable alternative to human evaluation, enabling large language models (LLMs) to provide reward signals in trainings. While recent work has explored multi-agent extensions such as multi-agent debate and meta-judging to enhance evaluation quality, the question of how intrinsic biases manifest in these settings remains underexplored. In this study, we conduct a systematic analysis of four diverse bias types: position bias, verbosity bias, chain-of-thought bias, and bandwagon bias. We evaluate these biases across two widely adopted multi-agent LLM-as-Judge frameworks: Multi-Agent-Debate and LLM-as-Meta-Judge. Our results show that debate framework amplifies biases sharply after the initial debate, and this increased bias is sustained in subsequent rounds, while meta-judge approaches exhibit greater resistance. We further investigate the incorporation of PINE, a leading single-agent debiasing method, as a bias-free agent within these systems. The results reveal that this bias-free agent effectively reduces biases in debate settings but provides less benefit in meta-judge scenarios. Our work provides a comprehensive study of bias behavior in multi-agent LLM-as-Judge systems and highlights the need for targeted bias mitigation strategies in collaborative evaluation settings.

大模型评判偏见分析多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。