研究大模型在开放式回答中如何受他人影响,发现评价者不中立且错误意见会破坏结果质量。
The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

- 设计新实验框架,分离出不同影响因素对模型修订的影响
- 所有模型在面对错误同伴答案时,修订质量均下降至最低
- 评价者主观倾向显著,需校准锚点避免误判
以往关于大模型从众性的研究多关注可验证标签下的离散答案切换。但开放式修订的答案质量是连续、隐含且评价不完美的,需采用新测量策略。本文提出一种跨混合主语料库与独立分解语料库的实验协议,可分离普通重答、候选内容暴露、同伴呈现残余效应及评价者对可见同伴上下文的方向性敏感。在四个开放权重生成器和三个基准测试中,所有错误同伴输入均导致各生成器-数据集组合中修订质量最低。对相同答案的盲评与明评存在差异:一名评价者向同伴观点靠拢,两人偏离,一人基本中立;GPT-4o 和 GPT-5.4-mini 的审计也非中立。此外,锚点审计显示,简洁正确锚点常被误读,若不显式校准将破坏隐含评分体系。结论为:翻转率不足以衡量开放式从众性,错误同伴有害,评价者非中立,锚点校准必要。
原文摘要 · Abstract (English)
Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strategy because answer quality is graded, latent, and judged imperfectly. We introduce an experimental protocol implemented across a pooled main peer-condition corpus and separately constructed decomposition corpora, allowing us to separate ordinary re-answering, candidate-content exposure, a bundled peer-presentation residual, and directional judge sensitivity to visible peer context. Across four open-weight generators and three benchmarks, all-wrong peer input produces the lowest-quality revisions in every generator-dataset cell. Blind and informed ratings of identical answers also differ by evaluator: one judge shifts toward the peer-endorsed position, two shift away, one is approximately neutral, and GPT-4o and GPT-5.4-mini audits are likewise non-neutral. Finally, an anchor audit shows that terse correct anchors can be misread often enough to destabilize the latent scale unless calibration is checked explicitly. These results support four conclusions: flip rates are insufficient as a complete measure of open-ended conformity, wrong peers harm open-ended revision, evaluators are not neutral, and anchor calibration is necessary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。