测试大模型理性与人类判断的匹配度,发现思考能提升理性但也更易受情绪影响。
Sparks of Rationality: Do Reasoning LLMs Align with Human Judgment and Choice?
- 通过思维链提升模型决策理性,逼近期望值最大化
- 情绪引导使模型行为出现方向性偏差,不同方法效果各异
- 适合研究人机决策差异或设计安全决策系统的人参考
大型语言模型(LLMs)正被用于招聘、医疗和经济判断等高风险决策场景,但真实人类判断兼具理性推理与情感偏见。为评估模型是否具备类似的人类非理性特征,我们测试了多个 LLM 家族在(i)理性选择核心公理基准,以及(ii)行为经济学和社交规范经典决策领域中的表现。结果显示,主动“思考”显著提升理性,推动模型趋向期望值最大化。为进一步探究情感干扰及其与推理的交互作用,我们采用两种情绪引导方法:上下文提示(ICP)和表示层引导(RLS)。ICP 引发强烈且难以校准的方向性偏移,而 RLS 产生更心理合理的模式但可靠性较低。结果表明,提升理性的机制同时放大对情绪干预的敏感性,不同引导方法在可控性与人类对齐之间存在权衡。整体揭示了推理与情感引导间的张力,对人类行为建模及基于 LLM 决策系统的安全部署具有启示。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly positioned as decision engines for hiring, healthcare, and economic judgment, yet real-world human judgment reflects a balance between rational deliberation and emotion-driven bias. If LLMs are to participate in high-stakes decisions or serve as models of human behavior, it is critical to assess whether they exhibit analogous patterns of (ir)rationalities and biases. To this end, we evaluate multiple LLM families on (i) benchmarks testing core axioms of rational choice and (ii) classic decision domains from behavioral economics and social norms where emotions are known to shape judgment and choice. Across settings, we show that deliberate "thinking" reliably improves rationality and pushes models toward expected-value maximization. To probe human-like affective distortions and their interaction with reasoning, we use two emotion-steering methods: in-context priming (ICP) and representation-level steering (RLS). ICP induces strong directional shifts that are often extreme and difficult to calibrate, whereas RLS produces more psychologically plausible patterns but with lower reliability. Our results suggest that the same mechanisms that improve rationality also amplify sensitivity to affective interventions, and that different steering methods trade off controllability against human-aligned behavior. Overall, this points to a tension between reasoning and affective steering, with implications for both human simulation and the safe deployment of LLM-based decision systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。