构建1.5万条因果常识问答数据集,评估大模型对实体因果推理能力。
CommonWhy: A Dataset for Evaluating Entity-Based Causal Commonsense Reasoning in Large Language Models

- 设计1.5万条为何类问题,聚焦实体因果推理
- 大模型在因果推理中频繁出现事实幻觉
- 适合评估大模型常识推理与知识图谱问答能力
为有效与现实世界交互,大语言模型需具备基于实体的常识推理能力,这要求将特定实体的事实知识与常识推断相结合。现有评估数据集多集中于真假判断或选择题,未能充分考察模型在因果关系上的溯因推理及解释生成能力。本文提出CommonWhy,包含1.5万条‘为何’类问题,用于评估大模型在实体因果关系上的常识推理能力。该数据集同时可作为知识图谱问答(KGQA)基准,所有答案所需知识均来自Wikidata知识图谱。与传统侧重事实检索的KGQA数据集不同,CommonWhy聚焦因果常识推理,确立了新型评估范式。实验表明,当前先进大模型及基于大模型的KGQA方法存在显著缺陷,包括频繁的事实幻觉和因果推理失败。
原文摘要 · Abstract (English)
To effectively interact with the real world, Large Language Models (LLMs) require entity-based commonsense reasoning, a challenging task that necessitates integrating factual knowledge about specific entities with commonsense inference. Existing datasets for evaluating LLM entity-based commonsense reasoning have largely focused on True/False or multiple-choice questions, leaving the explicit assessment of the model's ability in abductive reasoning about causes and effects and generating explanations largely unexamined. In this work, we introduce CommonWhy, a dataset of 15,000 why questions designed to evaluate entity-based commonsense reasoning about causal relationships in LLMs. CommonWhy also serves as a Knowledge Graph Question Answering (KGQA) benchmark, as all supporting knowledge required to answer its queries is available in the Wikidata knowledge graph. Unlike existing KGQA datasets, which primarily test fact retrieval, CommonWhy targets causal commonsense reasoning, establishing a new paradigm for KGQA evaluation. Experiments with state-of-the-art LLMs and LLM-based KGQA methods reveal their significant shortcomings, including frequent factual hallucinations and failures in causal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。