arXiv:2502.15487cs.CLcs.AI2025-02ACL被引 13

评测大模型在显式因果推理中的表现,发现顶级模型准确率仍不足0.8。

ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models

  • 构建包含因果与时间关系的多语言顺序数据集ExpliCa
  • 顶尖模型在因果推理上准确率未超0.8,易混淆时间与因果关系
  • 提示词与困惑度评分受模型规模影响不同,适合评估推理能力

大型语言模型(LLMs)在需要解释与推断准确性的任务中日益重要。本文提出ExpliCa,一个用于评估LLMs在显式因果推理方面表现的新数据集。ExpliCa独特地融合了以不同语言顺序呈现的因果与时间关系,并通过语言连接词明确表达。数据集还包含众包的人类可接受性评分。我们通过提示和基于困惑度的指标测试了七种商用与开源的LLMs,结果显示即使顶级模型在该任务上的准确率也未达0.80。有趣的是,模型常将时间关系误判为因果关系,且其表现受事件语言顺序显著影响。此外,困惑度得分与提示性能对模型规模的响应方式不同。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used in tasks requiring interpretive and inferential accuracy. In this paper, we introduce ExpliCa, a new dataset for evaluating LLMs in explicit causal reasoning. ExpliCa uniquely integrates both causal and temporal relations presented in different linguistic orders and explicitly expressed by linguistic connectives. The dataset is enriched with crowdsourced human acceptability ratings. We tested LLMs on ExpliCa through prompting and perplexity-based metrics. We assessed seven commercial and open-source LLMs, revealing that even top models struggle to reach 0.80 accuracy. Interestingly, models tend to confound temporal relations with causal ones, and their performance is also strongly influenced by the linguistic order of the events. Finally, perplexity-based scores and prompting performance are differently affected by model size.

因果推理大模型评估语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。