发现大模型多数“从众”其实与说话人无关,只是重复答案导致的。
Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks

- 通过移除说话人信息,发现重复错误答案仍会引发模型错误修正。
- 无说话人条件下66.5%正确答案被错误修改,远高于普通重问的10.3%。
- 提醒评估时需先测去说话人后的基础错误率,避免误判社会影响。
大模型的从众行为常表现为改变正确答案以趋同他人。我们发现,即使移除说话人,这种看似从众的现象仍广泛存在。原因是现有测评范式同时引入说话人存在和重复错误答案两个线索,难以区分真实社会影响。本文提出“无源条件”:仅保留重复答案但明确移除说话人。在六种开源大模型与七个人工智能问答及推理数据集上,该条件下初始正确的案例中有66.5%被错误修正,而普通重问仅10.3%。该效应在答案改写、选项隐藏等场景下依然显著。专家评议框架会提高此基线,但简单人物标签无效。模型出错时通常自信且无法通过简单校准恢复原答案。结论是:衡量社交影响应以去说话人后的基础错误率为基准,否则可能将重复文本误作社会压力。
原文摘要 · Abstract (English)
LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of this apparent conformity survives even after the peer is removed. The reason is a confound: standard conformity prompts mix two cues at once, the presence of a speaker and the repeated wrong answer itself. Existing benchmarks vary these cues together, so they cannot tell how much of the revision actually depends on the speaker. We introduce a no-source condition: the same asserted answer with the explicit speaker removed. Across six open-weight LLMs and seven QA and reasoning datasets, this condition alone causes harmful revision in $66.5\%$ of initially correct cases, compared with $10.3\%$ under a plain re-ask. The effect also remains when the repeated answer is paraphrased and when answer options are hidden in an open-ended setting. Source framing mainly modulates this floor: expert-panel framing raises it, while minimal person labels do not reliably raise it. When models flip, they are usually confidently wrong, and simple recalibration does not recover the original answer. Source attribution still matters, but it should be measured as an increment above this speaker-free floor. The methodological lesson is that conformity benchmarks should first measure what remains after the speaker is removed; without this step, benchmarks may mistake repeated text for social influence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。