arXiv:2410.11385cs.CL2024-10被引 3

测试大模型在未知因果场景下的推理能力,发现其表现受任务类型和术语干扰影响。

Do LLMs Have the Generalization Ability in Conducting Causal Inference?

  • 用随机生成的因果图构造新场景问题,评估模型泛化能力
  • 简单任务表现良好,但后门调整任务效果差且波动大
  • 熟悉术语即使新命名也会干扰模型,揭示认知偏差

在因果推断中,泛化能力指模型在新数据上对未知现象的因果效应进行估计的能力,对拓展知识边界至关重要。现有研究评估了大语言模型(LLMs)对已知现象的因果推断能力,但对未见现象的泛化能力尚未探索。本文选取因果路径发现(CP)、后门调整(BA)、事实推理(FI)和反事实推理(CI)四类代表性任务,提出一种基准生成框架:通过随机生成因果图与节点名,在假设的新因果情景中构建问题。基于此框架,构建了多复杂度的基准数据集,并在五种领先LLMs上系统测试其泛化性能。实验表明,尽管模型在简单CP、FI及复杂CI任务上表现较好,但在处理BA任务时困难重重,且性能随问题复杂度变化明显波动。此外,当现象名称包含已有术语(即使整体为新词),仍会因熟悉术语干扰导致泛化能力下降。

原文摘要 · Abstract (English)

In causal inference, generalization capability refers to the ability to conduct causal inference methods on new data to estimate the causal-effect between unknown phenomenon, which is crucial for expanding the boundaries of knowledge. Studies have evaluated the causal inference capabilities of Large Language Models (LLMs) concerning known phenomena, yet the generalization capabilities of LLMs concerning unseen phenomena remain unexplored. In this paper, we selected four tasks: Causal Path Discovery (CP), Backdoor Adjustment (BA), Factual Inference (FI), and Counterfactual Inference (CI) as representatives of causal inference tasks. To generate evaluation questions about previously unseen phenomena in new data on the four tasks, we propose a benchmark generation framework, which employs randomly generated graphs and node names to formulate questions within hypothetical new causal scenarios. Based on this framework, we compile a benchmark dataset of varying levels of question complexity. We extensively tested the generalization capabilities of five leading LLMs across four tasks. Experiment results reveal that while LLMs exhibit good generalization performance in solving simple CP, FI, and complex CI questions, they encounter difficulties when tackling BA questions and face obvious performance fluctuations as the problem complexity changes. Furthermore, when the names of phenomena incorporate existing terms, even if these names are entirely novel, their generalization performance can still be hindered by interference from familiar terms.

因果推断大模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。