arXiv:2505.24710cs.LGcs.AI2025-05IJCAI被引 7

让大模型学会因果推理,提升复杂环境下的决策能力。

Causal-aware Large Language Models: Enhancing Decision-Making Through Learning, Adapting and Acting

  • 引入因果模型构建环境结构知识,分学习-适应-行动三阶段迭代优化。
  • 在22个开放世界任务中,决策准确率显著高于传统大模型。
  • 适合需要动态适应和可靠推理的智能系统研发人员。

大语言模型虽具备丰富知识,但在决策中缺乏推理能力且难以适应新环境。为此,受人类认知启发,本文提出因果感知大模型(Causal-aware LLMs),将结构因果模型(SCM)融入决策流程,构建“学习-适应-行动”范式。学习阶段,利用大模型提取环境特定的因果实体与关系,初始化因果模型;适应阶段,通过外部反馈对因果模型进行因果干预更新;行动阶段,借助强化学习代理利用结构化因果知识高效制定策略。该过程迭代进行,使模型逐步掌握环境因果机制,实现更精准的理解与决策。在开放世界游戏Crafter中的22项多样化任务实验验证了方法的有效性。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown great potential in decision-making due to the vast amount of knowledge stored within the models. However, these pre-trained models are prone to lack reasoning abilities and are difficult to adapt to new environments, further hindering their application to complex real-world tasks. To address these challenges, inspired by the human cognitive process, we propose Causal-aware LLMs, which integrate the structural causal model (SCM) into the decision-making process to model, update, and utilize structured knowledge of the environment in a ``learning-adapting-acting" paradigm. Specifically, in the learning stage, we first utilize an LLM to extract the environment-specific causal entities and their causal relations to initialize a structured causal model of the environment. Subsequently,in the adapting stage, we update the structured causal model through external feedback about the environment, via an idea of causal intervention. Finally, in the acting stage, Causal-aware LLMs exploit structured causal knowledge for more efficient policy-making through the reinforcement learning agent. The above processes are performed iteratively to learn causal knowledge, ultimately enabling the causal-aware LLMs to achieve a more accurate understanding of the environment and make more efficient decisions. Experimental results across 22 diverse tasks within the open-world game ``Crafter" validate the effectiveness of our proposed method.

因果推理决策系统大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。