arXiv:2604.10511cs.AIcs.CL2026-04被引 1

LLM在政策评估中易受直觉影响,越反常识越容易出错。

Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation

论文配图:Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
图 1 · 摘自论文原文
  • 用思维链提示提升推理,但对反直觉问题效果大幅下降
  • 67.1%的判断差异源于问题是否符合常识,远超模型或提示方式
  • 模型知道答案却不会用,知识与推理能力不匹配

大型语言模型在因果与反事实推理中的应用日益广泛,但其在真实政策评估中的可靠性仍待考察。我们构建了一个包含40个实证政策评估案例的基准数据集,案例源自经济学与社会科学领域,基于同行评审证据,并按直觉性分类:符合常识(明显)、与常识模糊相关(模糊)或违背常识(反直觉)。我们在四个前沿大模型上测试五种提示策略,共进行8,000次实验,使用混合效应逻辑回归分析结果。发现三个关键点:(1) 思维链(CoT)悖论——思维链提示在明显案例中显著提升表现,但在反直觉案例中效果大幅减弱(交互作用比值比=0.278,p<0.001);(2) 直觉性是主导因素,案例级方差超过模型选择或提示策略(组内相关系数ICC=0.671);(3) 知识与推理分离——引用熟悉度与准确率无关(p=0.84),表明模型虽具备相关知识,却无法在结论违背直觉时有效运用。研究基于双过程理论(系统1与系统2)解释,认为当前大模型的“慢思考”仅形式上抑制直觉偏见,未真正实现深度推理。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for causal and counterfactual reasoning, yet their reliability in real-world policy evaluation remains underexplored. We construct a benchmark of 40 empirical policy evaluation cases drawn from economics and social science, each grounded in peer-reviewed evidence and classified by intuitiveness -- whether the empirical finding aligns with (obvious), is unclear relative to (ambiguous), or contradicts (counter-intuitive) common prior expectations. We evaluate four frontier LLMs across five prompting strategies with 8,000 experimental trials and analyze the results using mixed-effects logistic regression. Our findings reveal three key results: (1) a chain-of-thought (CoT) paradox, where chain-of-thought prompting dramatically improves performance on obvious cases but this benefit is substantially attenuated on counter-intuitive ones (interaction OR = 0.278, $p < 0.001$); (2) intuitiveness as the dominant factor, with case-level variance exceeding that of model choice or prompting strategy (ICC = 0.671); and (3) a knowledge-reasoning dissociation, where citation-based familiarity is unrelated to accuracy ($p = 0.84$), suggesting models possess relevant knowledge but fail to reason with it when findings contradict intuition. We frame these results through the lens of dual-process theory (System 1 vs. System 2) and argue that current LLMs' "slow thinking" achieves only partial inhibition of intuitive priors -- producing the form of deliberative reasoning without fully delivering its substance.

大模型推理反事实推理政策评估认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。