arXiv:2410.23884cs.LGcs.CL2024-10被引 15

分析大模型在叙事因果推理中的失效模式,发现其依赖表面规律而非深层逻辑。

Failure Modes of LLMs for Causal Reasoning on Narratives

  • 通过合成与真实数据测试模型因果推理能力
  • 发现模型常误用事件顺序或记忆知识而非上下文判断因果
  • 简单重述任务可显著提升模型推理鲁棒性

准确识别因果关系对自主决策和适应新场景至关重要。然而,精确推断因果结构需要结合世界知识与抽象逻辑推理。本文通过代表性叙事因果推理任务,研究这两者间的交互作用。在控制变量的合成、半合成及真实世界实验中,我们发现当前最先进的大语言模型(LLMs)常依赖表面启发式规则——例如仅根据事件先后顺序或回忆记忆中的常识来推断因果,而忽视上下文信息。此外,我们表明任务的简单改写能激发更稳健的推理行为。评估覆盖从线性链到包含碰撞器和分叉的复杂图结构等多种因果形式。这些发现揭示了大模型在因果推理中的系统性缺陷,为发展更符合原则性因果推断的方法提供了基础。

原文摘要 · Abstract (English)

The ability to robustly identify causal relationships is essential for autonomous decision-making and adaptation to novel scenarios. However, accurately inferring causal structure requires integrating both world knowledge and abstract logical reasoning. In this work, we investigate the interaction between these two capabilities through the representative task of causal reasoning over narratives. Through controlled synthetic, semi-synthetic, and real-world experiments, we find that state-of-the-art large language models (LLMs) often rely on superficial heuristics -- for example, inferring causality from event order or recalling memorized world knowledge without attending to context. Furthermore, we show that simple reformulations of the task can elicit more robust reasoning behavior. Our evaluation spans a range of causal structures, from linear chains to complex graphs involving colliders and forks. These findings uncover systematic patterns in how LLMs perform causal reasoning and lay the groundwork for developing methods that better align LLM behavior with principled causal inference.

因果推理大模型缺陷认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。