对比大模型即时与思考模式,发现推理能减少道德判断分歧。
How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models

- 在相同模型上对比即时与思考模式的道德判断差异。
- 推理使跨模型分歧平均提升至6.7/10,降低不一致性。
- 适合关注AI伦理一致性与模型行为可解释性的研究者。
我们评估了启用推理模式是否改变同一模型检查点内的道德判断。在100个道德判断场景和五个前沿推理训练大模型(Claude Sonnet 4.6、GPT 5.5、Gemini 3 Flash、DeepSeek V3.1、Qwen3.5 397B)上,即时模式与思考模式的整体二元判断一致性均保持高位,且统计上无显著差异(Krippendorff's alpha = 0.78 vs. 0.79)。然而,在21个存在模型争议的场景中,即时模式的一致性接近随机水平(alpha = 0.08)。在这些场景中,推理方向性地缩小了跨模型分歧,使平均成对一致性从5.4提升至6.7(满分10)。推理还降低了其中三个模型的群体判断不一致,未增加任何模型的不一致性。在所有五个模型家族中,推理比二元判断更频繁地改变自我标注的伦理框架。
原文摘要 · Abstract (English)
We evaluate whether enabling provider-exposed reasoning mode changes moral judgments within the same model checkpoint. Across 100 moral-judgment scenarios and five frontier reasoning-trained LLMs (Claude Sonnet 4.6, GPT 5.5, Gemini 3 Flash, DeepSeek V3.1, and Qwen3.5 397B), aggregate binary-verdict agreement remains high and statistically indistinguishable between instant and thinking modes (Krippendorff's alpha = 0.78 vs. 0.79). However, disagreement is concentrated in 21 model-disputed scenarios, where instant-mode agreement is near chance (alpha = 0.08). On these scenarios, reasoning directionally narrows cross-model disagreement, increasing mean pairwise agreement from 5.4 to 6.7 out of 10. Reasoning also reduces demographic-judgment inconsistency in three of five models and does not increase it for any model. Across all five model families, reasoning changes self-labeled ethical frameworks more often than binary verdicts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。