大模型在因果判断中易产生虚假因果错觉,可能误导重要决策。
Do Large Language Models Show Biases in Causal Learning? Insights from Contingency Judgment
- 用1000个无因果关系的医学场景测试大模型判断能力
- 所有模型均错误判断出不存在的因果关系,存在系统性偏差
- 提示研究者警惕模型在医疗等领域的误判风险
因果学习是基于现有信息进行因果推断的认知过程,常受认知偏差影响。其中,虚假因果幻觉指在缺乏证据时仍感知变量间存在因果关系,可能引发社会偏见、刻板印象、虚假信息传播等问题。本文通过经典认知科学范式——条件判断任务,检验大语言模型是否易出现此类偏差。我们构建了1000个医学情境下的零因果关系数据集,要求模型评估潜在原因的有效性。结果发现,所有被测模型均系统性地推断出不存在的因果关系,表现出强烈的虚假因果倾向。尽管学界对大模型是否真正理解因果关系仍有争议,但本研究支持其仅能复现因果语言而无深层理解的观点,警示其在需精准因果推理的领域(如医疗)应用可能带来严重风险。
原文摘要 · Abstract (English)
Causal learning is the cognitive process of developing the capability of making causal inferences based on available information, often guided by normative principles. This process is prone to errors and biases, such as the illusion of causality, in which people perceive a causal relationship between two variables despite lacking supporting evidence. This cognitive bias has been proposed to underlie many societal problems, including social prejudice, stereotype formation, misinformation, and superstitious thinking. In this work, we examine whether large language models are prone to developing causal illusions when faced with a classic cognitive science paradigm: the contingency judgment task. To investigate this, we constructed a dataset of 1,000 null contingency scenarios (in which the available information is not sufficient to establish a causal relationship between variables) within medical contexts and prompted LLMs to evaluate the effectiveness of potential causes. Our findings show that all evaluated models systematically inferred unwarranted causal relationships, revealing a strong susceptibility to the illusion of causality. While there is ongoing debate about whether LLMs genuinely understand causality or merely reproduce causal language without true comprehension, our findings support the latter hypothesis and raise concerns about the use of language models in domains where accurate causal reasoning is essential for informed decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。