测试大模型是否像人一样有因果偏见,发现它们更依赖规则而非人类的直觉推理。
Do LLMs Share Human-Like Biases? Causal Reasoning Under Prior Knowledge, Irrelevant Context, and Varying Compute Budgets
- 用碰撞器结构任务对比20多个大模型与人类的因果判断
- 多数模型不表现出人类典型的弱解释消除和马尔可夫违反偏见
- 思维链提示能提升模型在干扰下的推理鲁棒性,适合安全部署研究
大型语言模型(LLMs)在需要因果推理的领域应用日益广泛,但其判断是否反映规范的因果计算、类似人类的捷径,或脆弱的模式匹配仍不明确。我们在11个由碰撞器结构($C_1 \rightarrow E \leftarrow C_2$)形式化的因果判断任务上,对20多个LLMs与匹配的人类基准进行评估。结果发现,一个小型可解释模型能很好地压缩LLMs的因果判断;多数LLMs采用更规则化的推理策略,而人类则会考虑未提及的潜在因素。此外,大多数LLMs并未表现出人类特有的弱解释消除和马尔可夫违反偏见。我们进一步测试了模型在语义抽象和提示过载(注入无关文本)下的鲁棒性,发现思维链(CoT)可增强多数模型的稳定性。该差异表明,当人类偏见不可取时,LLMs可作为补充,但其规则化推理在内在不确定性下可能失效,凸显了为安全有效部署而刻画其推理策略的重要性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in domains where causal reasoning matters, yet it remains unclear whether their judgments reflect normative causal computation, human-like shortcuts, or brittle pattern matching. We benchmark 20+ LLMs against a matched human baseline on 11 causal judgment tasks formalized by a collider structure ($C_1 \rightarrow E \leftarrow C_2$). We find that a small interpretable model compresses LLMs' causal judgments well and that most LLMs exhibit more rule-like reasoning strategies than humans who seem to account for unmentioned latent factors in their probability judgments. Furthermore, most LLMs do not mirror the characteristic human collider biases of weak explaining away and Markov violations. We probe LLMs' causal judgment robustness under (i) semantic abstraction and (ii) prompt overloading (injecting irrelevant text), and find that chain-of-thought (CoT) increases robustness for many LLMs. Together, this divergence suggests LLMs can complement humans when known biases are undesirable, but their rule-like reasoning may break down when uncertainty is intrinsic - highlighting the need to characterize LLM reasoning strategies for safe, effective deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。