测试大模型能否分清因果解释与语义相关,发现它们常混淆两者。
CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models
- 构建3000个陈述-理由对,区分语义关联与真实因果关系
- 最大模型准确率仅0.55(马修斯相关系数),提升有限
- 参数越大越容易误判,从过度怀疑转为过度接受因果
我们提出CLEAR-3K,一个包含3,000个断言-推理问题的数据集,用于评估语言模型是否能判断一个陈述是否在因果上解释另一个。每个问题提供一对陈述,要求模型区分语义相关性和真正的因果解释关系。通过对21个前沿语言模型(参数量从0.5B到72B)的全面评估,我们发现两个核心问题:第一,模型常将语义相似性误认为因果关系,依赖词汇和语义重叠而非推断真实因果;第二,随着参数规模增加,模型从过度怀疑因果关系转向过度宽容地接受因果,但即便最优模型,其马修斯相关系数(MCC)仍停滞在0.55。因此,CLEAR-3K为发展和评估语言模型的真实因果推理能力提供了关键基准,这对需要准确因果评估的应用至关重要。
原文摘要 · Abstract (English)
We introduce CLEAR-3K, a dataset of 3,000 assertion-reasoning questions designed to evaluate whether language models can determine if one statement causally explains another. Each question present an assertion-reason pair and challenge language models to distinguish between semantic relatedness and genuine causal explanatory relationships. Through comprehensive evaluation of 21 state-of-the-art language models (ranging from 0.5B to 72B parameters), we identify two fundamental findings. First, language models frequently confuse semantic similarity with causality, relying on lexical and semantic overlap instead of inferring actual causal explanatory relationships. Second, as parameter size increases, models tend to shift from being overly skeptical about causal relationships to being excessively permissive in accepting them. Despite this shift, performance measured by the Matthews Correlation Coefficient plateaus at just 0.55, even for the best-performing models.Hence, CLEAR-3K provides a crucial benchmark for developing and evaluating genuine causal reasoning in language models, which is an essential capability for applications that require accurate assessment of causal relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。