研究发现模型身份偏见在多智能体评估中被部分匿名化掩盖,需全管道匿名才能真实检测。
Peer Identity Bias in Multi-Agent LLM Evaluation: An Empirical Study Using the TRUST Democratic Discourse Analysis Pipeline
- 通过多通道暴露模型身份,发现单通道匿名化会因正负抵消而误判无偏见
- 完全匿名下,同质模型组放大逢迎倾向,异质组则相反
- 异质模型组合更稳健,且需全管道匿名才可有效验证系统可靠性
TRUST民主话语分析流水线通过多个结构通道让大语言模型组件暴露于同行模型身份,但其偏见影响此前未被实证检验。本文首次系统测量了TRUST中所有活跃身份暴露通道的依赖性评分偏见,覆盖四个模型家族、两种匿名范围,针对30个政治声明进行测试。核心发现为:单通道匿名化几乎消除偏见效应,因各通道作用方向相反相互抵消——这会导致评估者误判身份偏见不存在。唯有全流水线匿名才能揭示真实模式:同质模型集合在身份可见时加剧身份驱动的逢迎行为,而异质生产配置则呈现相反趋势。模型选择独立影响结果:某一测试模型的基础逢迎程度是其他模型的2至3倍,且在意识形态议题上几乎无辩论冲突,使其在依赖角色间真实分歧的系统中结构上不适用。三条实践结论:第一,异质模型集合比同质集合结构更稳健,达成更高共识率并降低身份放大;第二,完整流水线匿名是有效偏见测量的必要条件,部分匿名不足且具误导性;第三,这些发现直接影响高质量关键应用中多智能体大模型系统的验证:在部分匿名或同质集合下通过验证的系统,可能仍存在单通道测量无法察觉的结构性身份偏见。
原文摘要 · Abstract (English)
The TRUST democratic discourse analysis pipeline exposes its large language model (LLM) components to peer model identity through multiple structural channels -- a design feature whose bias implications have not previously been empirically tested. We provide the first systematic measurement of identity-dependent scoring bias across all active identity exposure channels in TRUST, crossing four model families with two anonymization scopes across 30 political statements. The central finding is that single-channel anonymization produces near-zero bias effects, because individual channels act in opposite directions and cancel each other out -- a result that would lead an evaluator to conclude that identity bias is absent when it is not. Only full-pipeline anonymization reveals the true pattern: homogeneous ensembles amplify identity-driven sycophancy when model identity is fully visible, while the heterogeneous production configuration shows the reverse. Model choice matters independently: one tested model exhibits baseline sycophancy two to three times higher than the others and near-zero deliberative conflict on ideological topics, making it structurally unsuitable for pipelines where genuine inter-role disagreement is the intended quality mechanism. Three practical conclusions follow. First, heterogeneous model ensembles are structurally more robust than homogeneous ones, achieving higher consensus rates and lower identity amplification. Second, full-pipeline anonymization is required for valid bias measurement -- partial anonymization is insufficient and actively misleading. Third, these findings have direct implications for the validation of multi-agent LLM systems in quality-critical applications: a system validated under partial anonymization or with a homogeneous ensemble may pass validation while retaining structural identity bias invisible to single-channel measurement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。