LLM的'是/否'偏见源于答题顺序和用词,非道德判断变化。
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

- 通过交叉对称实验分离语义、顺序与词汇影响
- 前沿模型道德立场稳定,但部分模型受'否'字和末位选项影响大
- 适合研究模型决策机制或评估生成公平性的研究人员
大型语言模型(LLMs)在判断中常呈现二元结论,已有研究发现其判断会因无关语义的表述变化而改变,例如在道德困境中出现放大化的'是/否'偏见,人类则无此现象。单一提问方式无法区分这种偏移的本质:在'是/否'问题中,'否'同时是逻辑结论、词汇标记和最后显示选项。本文引入心理测量学实验范式——交叉对称化,即对所有逻辑无关因素进行平衡配对翻转,在一个包含多种题型的语料库中测试。对逻辑等价题型的分级评分恢复出一致的内在道德尺度:前沿模型的立场θ几乎不受格式影响(跨形式不一致性0.12–0.21,轴向±1);小规模开源模型则表现出特定模式的失效。强制使用'是/否'标签叠加了可分解的人工效应:对最后出现选项的顺序偏倚(与人类相反),以及对'否'字的词汇吸引;该人工效应仅在Claude模型中显著(故事平均-0.32至-0.86),GPT-5.5和Gemini接近0,且在长推理下减弱。当将词语与结论解耦(如用任意标签替换'是/否'),结论关联的逻辑偏倚对所有前沿模型均≈0,但模型特有的标签与顺序依赖仍存在:模型并非倾向拒绝,而是受表面打印内容驱动。一个最小模型 $P = σ((θ\pm m)/s)$ 可以用框架敏感度m和道德决断力s来量化任意人工效应,与采样温度明确区分。该方法适用于任何困境集和二元格式:衡量模型价值必须跨框架测试,而非单次提问。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word "no" is at once logical verdict, lexical token, and last-printed option. We introduce a psychometric battery that separates these: crossed symmetrization - every logically irrelevant factor flipped in balanced pairs - across a corpus of question forms. A graded rating across logically equivalent forms recovers a coherent internal moral scale: frontier models' stance $θ$ is nearly format-invariant (cross-form incoherence 0.12-0.21 on a $\pm 1$ axis); small open-weight models fail in model-specific ways. Forcing the verdict through yes/no overlays a decomposable artifact: an order bias toward the last-printed option - opposite to classic human primacy - plus a lexical pull toward the word "no"; the artifact is substantial only in the Claude models (story-averaged -0.32 to -0.86), $\approx 0$ for GPT-5.5 and Gemini, and shrinks under extended reasoning. The word and the verdict share one token; swapping the words for arbitrary labels separates them, and the verdict-attached logical bias proves $\approx 0$ for every frontier model, while model-specific label and order attachments remain: the models are not drawn toward rejecting - the pull follows the printed surface, not the verdict it carries. A minimal model, $P = σ((θ\pm m)/s)$, summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, measurably distinct from sampling temperature. The battery applies unchanged to any dilemma set and binary format: measuring what a model values requires crossing the frames of the question, not asking once.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。