构建首个基于图与表格数据的因果推理评测基准,揭示大模型在真实场景下的短板。
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
- 用图和表格数据设计多任务评测框架,覆盖因果推理全链条
- 开源模型在表格数据上发现新规律能力薄弱,准确率不足40%
- 揭示不同类别任务间相关性高于同类任务,为模型评估提供新视角
因果推理对大语言模型在教育、医疗等领域的应用至关重要,但现有基准多聚焦对话、数学和编程任务,难以评估真实世界问题解决能力。本文提出CARL-GT基准,通过图结构和表格数据评估大模型在因果图推理、知识发现和决策制定方面的表现。该基准涵盖多样化任务,并设计有效零样本提示。实验评估多个开源模型,发现其在表格数据中发现新洞察的能力普遍较弱(准确率低于40%)。进一步分析显示,不同任务类别间性能相关性高于同一类别内部,表明因果推理能力具有跨任务协同特征。
原文摘要 · Abstract (English)
Causal reasoning capabilities are essential for large language models (LLMs) in a wide range of applications, such as education and healthcare. But there is still a lack of benchmarks for a better understanding of such capabilities. Current LLM benchmarks are mainly based on conversational tasks, academic math tests, and coding tests. Such benchmarks evaluate LLMs in well-regularized settings, but they are limited in assessing the skills and abilities to solve real-world problems. In this work, we provide a benchmark, named by CARL-GT, which evaluates CAusal Reasoning capabilities of large Language models using Graphs and Tabular data. The benchmark has a diverse range of tasks for evaluating LLMs from causal graph reasoning, knowledge discovery, and decision-making aspects. In addition, effective zero-shot learning prompts are developed for the tasks. In our experiments, we leverage the benchmark for evaluating open-source LLMs and provide a detailed comparison of LLMs for causal reasoning abilities. We found that LLMs are still weak in casual reasoning, especially with tabular data to discover new insights. Furthermore, we investigate and discuss the relationships of different benchmark tasks by analyzing the performance of LLMs. The experimental results show that LLMs have different strength over different tasks and that their performance on tasks in different categories, i.e., causal graph reasoning, knowledge discovery, and decision-making, shows stronger correlation than tasks in the same category.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。