语言模型在因果推理中存在偏向简单关系的系统性偏差,影响决策可靠性。
Language Agents Mirror Human Causal Reasoning Biases. How Can We Help Them Think Like Scientists?
- 用儿童心理学实验测试模型因果探索能力
- 模型对常见因果关系准确,对复杂组合关系显著失败
- 提出采样修正方法可有效缓解偏差,适合需严谨推理的应用
语言模型代理作为自主决策者,需高效探索世界因果结构以实现稳健推理。本文采用发展心理学中的经典Blicket测试范式,考察模型对因果关系的探索与推断能力。结果发现,模型能可靠识别常见、直观的析取性因果关系,但系统性地难以处理不常见却证据充分的合取性因果关系,这种“析取性偏差”在不同模型家族、规模和提示策略下均存在,且任务复杂度越高表现越差。有趣的是,人类成人也表现出类似偏差,表明模型可能从训练数据中继承了深层认知启发式。我们量化了模型与人类的推理相似性,发现模型呈现成人而非儿童的推理模式。为此,我们提出一种测试时采样方法,主动筛选并排除因果假设,该可扩展方法显著降低了析取性偏差,推动模型向科学化、因果严谨的推理目标迈进。
原文摘要 · Abstract (English)
Language model (LM) agents are increasingly used as autonomous decision-makers which need to actively gather information to guide their decisions. A crucial cognitive skill for such agents is the efficient exploration and understanding of the causal structure of the world -- key to robust, scientifically grounded reasoning. Yet, it remains unclear whether LMs possess this capability or exhibit systematic biases leading to erroneous conclusions. In this work, we examine LMs' ability to explore and infer causal relationships, using the well-established Blicket Test paradigm from developmental psychology. We find that LMs reliably infer the common, intuitive disjunctive causal relationships but systematically struggle with the unusual, yet equally (or sometimes even more) evidenced conjunctive ones. This "disjunctive bias" persists across model families, sizes, and prompting strategies, and performance further declines as task complexity increases. Interestingly, an analogous bias appears in human adults, suggesting that LMs may have inherited deep-seated reasoning heuristics from their training data. To this end, we quantify similarities between LMs and humans, finding that LMs exhibit adult-like inference profiles (but not child-like). Finally, we propose a test-time sampling method which explicitly samples and eliminates hypotheses about causal relationships from the LM. This scalable approach significantly reduces the disjunctive bias and moves LMs closer to the goal of scientific, causally rigorous reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。