评估大模型因果推理能力,梳理提升方法并指明未来方向
CausalEval: Towards Better Causal Reasoning in Language Models
- 将模型分为推理引擎与辅助工具两类,系统梳理现有方法
- 实证测试发现当前模型在因果任务中表现仍不理想
- 适合关注大模型认知能力与可信AI的研究者参考
因果推理是智能的核心,对问题解决、决策和世界理解至关重要。尽管语言模型能生成推理过程,但其可靠进行因果推理的能力仍不确定,尤其在需要深层因果理解的任务中表现不佳。本文提出CausalEval,全面回顾旨在提升语言模型因果推理能力的研究,并对当前模型与方法进行实证评估。我们根据语言模型的作用将其分为两类:作为推理引擎或作为传统因果推理方法的知识/数据提供者,并深入讨论各类别中的方法论。随后,我们在一系列因果推理任务上评估当前模型及增强方法的表现,给出关键发现与深入分析。最后,总结现有研究洞见,指出未来有潜力的研究方向。本工作旨在成为该领域的综合性资源,推动语言模型因果推理的进一步发展。
原文摘要 · Abstract (English)
Causal reasoning (CR) is a crucial aspect of intelligence, essential for problem-solving, decision-making, and understanding the world. While language models (LMs) can generate rationales for their outputs, their ability to reliably perform causal reasoning remains uncertain, often falling short in tasks requiring a deep understanding of causality. In this paper, we introduce CausalEval, a comprehensive review of research aimed at enhancing LMs for causal reasoning, coupled with an empirical evaluation of current models and methods. We categorize existing methods based on the role of LMs: either as reasoning engines or as helpers providing knowledge or data to traditional CR methods, followed by a detailed discussion of methodologies in each category. We then assess the performance of current LMs and various enhancement methods on a range of causal reasoning tasks, providing key findings and in-depth analysis. Finally, we present insights from current studies and highlight promising directions for future research. We aim for this work to serve as a comprehensive resource, fostering further advancements in causal reasoning with LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。