arXiv:2608.04415cs.CL2026-08

大模型安全评审团易受同伴误导,导致误判率飙升至100%。

Social Pressure Breaks Majority Voting in LLM Safety Panels

  • 让多个大模型先独立判断,再受模拟同伴影响后重判。
  • 同伴错误引导使误报率从56.5%升至87.5%,集体投票达100%。
  • 模型更易受'不安全'信号影响,适合关注AI安全评估的研究者。

大型语言模型常用于检测不安全内容,常见做法是通过多模型评审团进行多数决以纠正个体错误。然而,当所有模型在投票前看到相同的误导性上下文时,这一优势可能失效。我们设计了一个两轮控制实验:每个模型先独立判断一项内容,再在六名模拟同伴或宣称错误标签或选择沉默后重新判断。最终以多数票决定结果。在六种开源大模型和六个数据集上,发现若同伴给出错误标签,平均误报率从沉默状态下的56.5%上升至87.5%,而多数决使评审团整体误报率高达100%。若无标签提示,该评审团表现优于单个模型。效果高度不对称:模型更易响应‘不安全’引导(约75%),对‘安全’引导反应微弱(约17%),导致误报率急剧上升,但漏检率基本不变。专有模型测试显示不同模型间差异显著。研究揭示了共享社会线索是安全评审团的失效模式,并提出简单预部署诊断方法。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but this benefit may disappear when every model sees the same misleading context before voting. We study this risk in a controlled two-round experiment. Each model first judges an item alone, then judges it again after six simulated peers either assert the wrong label or abstain. We combine the final judgments by majority vote. Across six open-weight LLMs and six datasets, we find that the wrong-label peer message raises the average reviewer false-alarm rate from 56.5% under silent peers to 87.5%, and majority voting raises the panel false-alarm rate to 100%. Without an asserted label, the same panel outperforms its average member. The effect is strongly asymmetric: reviewers follow pushes toward "unsafe" far more than pushes toward "safe" (about 75% versus 17%), so the panel's false-alarm rate rises sharply while its harmful-miss rate changes little. The proprietary-model probe shows substantial variation across models. These results identify susceptibility to shared social cues as a failure mode of safety panels and provide a simple pre-deployment diagnostic.

大模型安全群体决策误报率社会压力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。