arXiv:2608.30373cs.CL2026-08中稿 · EMNLP

多智能体辩论让评分更一致,却反而偏离人类判断。

Beyond Consensus: Downward Bias and Role Asymmetry in Multi-Agent LLM Judges for Subjective Evaluation

论文配图:Beyond Consensus: Downward Bias and Role Asymmetry in Multi-Agent LLM Judges for Subjective Evaluation
图 1 · 摘自论文原文
  • 用对称角色代替严格角色,避免评分偏倚。
  • 共识机制反而让评分偏离人类标准,平均相关性下降12%。
  • 适合关注主观评价公平性的评测系统设计者。

多智能体辩论(MAD)通过多个代理协商达成共识,被广泛用于提升大模型评估效果。然而,在基于主观评分标准的任务中,代理间的一致性并不等于与人类判断对齐。本文对比单裁判基线与基于共识的MAD协议在六种大模型上的主观评价表现,并设计三种消融实验,分离角色提示、多轮交互和显式分数共享的影响。结果表明,单裁判基线在六种裁判模型上平均人类对齐度最高,而MAD在两项任务中均出现人类对齐度下降。消融分析显示,性能下降主要源于角色不对称,而非交互本身。赋予严格评判角色导致系统性向下偏倚,且共识过程无法纠正该偏倚。关键发现是:共识并非简单平均,而是严格立场主导,最终得分远低于严格与宽松条件的算术中点。消除角色不对称(对称MAD)可显著恢复基线表现,而隐藏同行分数则加剧分歧并降低平均人类对齐度。这揭示了角色专业化共识协议在主观评分中的结构性缺陷。

原文摘要 · Abstract (English)

Multi-Agent Debate (MAD) has been widely adopted to improve LLM-based evaluation by prompting multiple agents to negotiate and reach a consensus. However, for subjective rubric-based scoring, inter-agent agreement does not guarantee alignment with human judgments. In this paper, we compare a single-judge baseline against a consensus-based MAD protocol on subjective evaluation tasks and design three ablations to isolate the impact of role prompting, multi-round interaction, and explicit score sharing. Evaluations across six LLMs show that the single-judge baseline achieves the strongest human alignment on average across six judge models, whereas MAD shows degradation in human alignment on both tasks. Our ablations demonstrate that this performance drop stems primarily from asymmetric role prompting rather than the interaction itself. Specifically, assigning a strict judge role introduces a systematic downward bias that the consensus process fails to correct. The central finding is that this bias reflects strict-stance dominance beyond averaging: the consensus score falls well beyond the arithmetic midpoint of the standalone strict and lenient conditions, rather than averaging them out. Removing role asymmetry (Symmetric MAD) largely recovers baseline performance, while masking peer scores widens inter-agent disagreement on average and worsens average human alignment. These findings demonstrate that multi-agent consensus can enforce artificial agreement at the expense of true human alignment, revealing a structural limitation in consensus-style, role-specialized MAD protocols for subjective scoring.

多智能体主观评价评分偏差一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。