用因果贝叶斯网络解析大模型推理,发现其与人类在因果思维上既相似又不同。
Causal Strengths and Leaky Beliefs: Interpreting LLM Reasoning via Noisy-OR Causal Bayes Nets
- 构建漏失型噪声或门贝叶斯网络,量化大模型的因果推理模式。
- 20+大模型在11个任务中表现不一,部分与人类推理一致但存在系统性偏差。
- 可识别模型与人类的推理特征差异,适合研究认知机制的学者参考。
人类与机器智能的本质是长期未解之谜。尽管尚无普遍接受的定义,但因果推理常被视为智能的核心特征(Lake et al., 2017)。在相同任务上评估大模型与人类的因果推理能力,有助于更全面理解两者的优劣。本研究提出三个问题:(Q1) 大模型是否与人类在相同推理任务上保持一致?(Q2) 大模型与人类在任务层面是否具有一致性?(Q3) 它们是否存在不同的推理特征?我们通过在十一组语义明确的因果任务上评估20多个大模型,任务结构基于碰撞图 $C_1 o E rom C_2$,采用直接响应(单次输出数值 = 查询节点为1的概率判断)和思维链(CoT;先思考,再作答)两种方式。推理判断通过一个包含参数 $θ=(b,m_1,m_2,p(C)) imes [0,1]$ 的漏失型噪声或门因果贝叶斯网络(CBN)建模,其中共享先验 $p(C)$。通过AIC比较三参数对称因果强度($m_1=m_2$)与四参数非对称变体($m_1≠m_2$)的拟合效果,选择最优模型。
原文摘要 · Abstract (English)
The nature of intelligence in both humans and machines is a longstanding question. While there is no universally accepted definition, the ability to reason causally is often regarded as a pivotal aspect of intelligence (Lake et al., 2017). Evaluating causal reasoning in LLMs and humans on the same tasks provides hence a more comprehensive understanding of their respective strengths and weaknesses. Our study asks: (Q1) Are LLMs aligned with humans given the \emph{same} reasoning tasks? (Q2) Do LLMs and humans reason consistently at the task level? (Q3) Do they have distinct reasoning signatures? We answer these by evaluating 20+ LLMs on eleven semantically meaningful causal tasks formalized by a collider graph ($C_1\!\to\!E\!\leftarrow\!C_2$ ) under \emph{Direct} (one-shot number as response = probability judgment of query node being one and \emph{Chain of Thought} (CoT; think first, then provide answer). Judgments are modeled with a leaky noisy-OR causal Bayes net (CBN) whose parameters $θ=(b,m_1,m_2,p(C)) \in [0,1]$ include a shared prior $p(C)$; we select the winning model via AIC between a 3-parameter symmetric causal strength ($m_1{=}m_2$) and 4-parameter asymmetric ($m_1{\neq}m_2$) variant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。