大模型因果推理比人类更准,但仍有盲区。
Do Large Language Models Reason Causally Like Us? Even Better?
- 用碰撞图任务测试人类与大模型的因果推断能力。
- GPT-4o等模型表现优于人类,且无关联偏差影响。
- 仍无法完全理解'解释消解'等复杂因果机制。
因果推理是智能的核心。大语言模型(LLMs)在生成类人文本方面表现出色,引发对其回答是否反映真实理解还是统计模式的疑问。我们通过基于碰撞图的任务,比较了人类与四种LLMs的因果推理能力,评估在给定其他变量证据下,查询变量出现的可能性。结果显示,LLMs的因果推断从常不合理(GPT-3.5)到类人,再到通常比人类更符合规范(GPT-4o、Gemini-Pro和Claude)。计算模型拟合表明,GPT-4o、Gemini-Pro和Claude表现优异的一个原因是它们未表现出困扰人类因果推理的“关联偏差”。然而,即使这些模型也未能完全捕捉与碰撞图相关的细微推理模式,如“解释消解”(explaining away)。
原文摘要 · Abstract (English)
Causal reasoning is a core component of intelligence. Large language models (LLMs) have shown impressive capabilities in generating human-like text, raising questions about whether their responses reflect true understanding or statistical patterns. We compared causal reasoning in humans and four LLMs using tasks based on collider graphs, rating the likelihood of a query variable occurring given evidence from other variables. LLMs' causal inferences ranged from often nonsensical (GPT-3.5) to human-like to often more normatively aligned than those of humans (GPT-4o, Gemini-Pro, and Claude). Computational model fitting showed that one reason for GPT-4o, Gemini-Pro, and Claude's superior performance is they didn't exhibit the "associative bias" that plagues human causal reasoning. Nevertheless, even these LLMs did not fully capture subtler reasoning patterns associated with collider graphs, such as "explaining away".
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。