arXiv:2411.19500cs.CLcs.AI2024-11NeurIPS被引 11

构建日常活动因果推理框架,评估大模型对真实世界事件的理解能力。

COLD: Causal reasOning in cLosed Daily activities

  • 基于人类对日常活动的常识,构建可生成上百万因果问题的推理框架。
  • 大模型在简单日常任务上仍难完成因果推理,表明其理解能力有限。
  • 通过反向门准则量化事件间因果强度,提供可验证的分析方法。

大型语言模型在算术和推理等任务中表现卓越,但要评估其认知能力,因果推理成为检验其对世界机制理解的重要指标。现有研究或聚焦开放域因果常识推理(CCR),缺乏理论验证;或依赖符号化表示的因果推断引擎,脱离真实场景。本文提出COLD(Causal reasOning in cLosed Daily activities)框架,基于人类对日常活动的理解,构建真实世界情境下的因果推理体系。该框架可生成约900万条因果查询,接近迷你图灵测试,用于评估模型对日常任务的因果理解。实验表明,即使对人类而言极为简单的日常活动,大模型也难以完成因果推理。进一步利用反向门准则分析事件间的因果强度,揭示模型在因果建模上的薄弱环节。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown state-of-the-art performance in a variety of tasks, including arithmetic and reasoning; however, to gauge the intellectual capabilities of LLMs, causal reasoning has become a reliable proxy for validating a general understanding of the mechanics and intricacies of the world similar to humans. Previous works in natural language processing (NLP) have either focused on open-ended causal reasoning via causal commonsense reasoning (CCR) or framed a symbolic representation-based question answering for theoretically backed-up analysis via a causal inference engine. The former adds an advantage of real-world grounding but lacks theoretically backed-up analysis/validation, whereas the latter is far from real-world grounding. In this work, we bridge this gap by proposing the COLD (Causal reasOning in cLosed Daily activities) framework, which is built upon human understanding of daily real-world activities to reason about the causal nature of events. We show that the proposed framework facilitates the creation of enormous causal queries (~ 9 million) and comes close to the mini-turing test, simulating causal reasoning to evaluate the understanding of a daily real-world task. We evaluate multiple LLMs on the created causal queries and find that causal reasoning is challenging even for activities trivial to humans. We further explore (the causal reasoning abilities of LLMs) using the backdoor criterion to determine the causal strength between events.

因果推理大模型评估常识理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。