评估大模型在因果边判断中的可靠性,发现其易过拟合且自信过度。
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

- 用多种提示策略测试12个开源大模型的因果边预测能力
- 40%间接边和36%反向边被误判为直接边,且80%错误判断信心超80%
- 模型间一致性比内部置信度更可信,适合做因果先验而非直接证据
大语言模型(LLMs)被越来越多地用于提供结构因果发现的先验因果知识,但其直接边判断与置信度是否可信仍不明确。我们系统评估了12个指令微调的开源模型,在六个基准因果图上,采用五种提示策略和四种置信度来源:口头表达、基于对数几率、跨提示一致性和跨模型一致性。在纯语言成对协议下,研究得出三个关键发现:(i) LLM因果判断强偏向召回率:模型预测出过于密集的图,存在大量假阳性边;提示策略仅改变精确率-召回率权衡,无法解决过预测问题。模型规模增益在最大图上减弱,且不能消除校准偏差。(ii) LLM常捕捉因果相关性,但难以可靠识别直接性或方向性。相对于已发布参考图,模型将40.0%的间接边和36.0%的反向非边误判为直接边,而其他非边误判率为28.2%。此外,80.8%和84.6%的此类错误判断获得至少80%的口头置信度,显示结构性错误下的严重过度自信。(iii) 传统置信度估计不可靠,而一致性信号更具前景。基于对数几率的置信度常坍缩至接近1.0,无论正确与否;而跨提示和跨模型一致性在平均校准和区分度上表现更好,尽管经霍尔姆校正后优势不显著。对基准熟悉性审计发现五个模型-数据集组合存在潜在熟悉性,均涉及AsiaM。总体而言,我们的结果表明,应将LLMs视为需外部验证的软因果先验,而非因果结构的直接证据。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources: verbalized, logit-based, cross-prompt agreement, and cross-model agreement. Under our language-only pairwise protocol, our evaluation yields three key findings. (i) LLM-based causal judgments are strongly recall-dominant: models predict overly dense graphs with many false-positive edges, while prompting mainly shifts the precision-recall trade-off rather than resolving overprediction. Gains from model scale diminish on the largest graphs and do not eliminate miscalibration. (ii) LLMs often capture causal relatedness without reliably identifying directness or orientation. Relative to published reference graphs, models misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges, versus 28.2% of other non-edges. Moreover, 80.8% and 84.6% of these false positives receive verbalized confidence of at least 80%, revealing substantial overconfidence in structurally incorrect predictions. (iii) Conventional confidence estimates are unreliable, whereas agreement offers a more promising signal. Logit-based confidence frequently collapses near 1.0 regardless of correctness, while cross-prompt and cross-model agreement achieve better mean calibration and discrimination, though their advantages are not statistically significant after Holm correction. A benchmark-familiarity audit further identifies potential familiarity in five model-dataset pairs, all involving AsiaM. Overall, our results suggest LLMs are better viewed as sources of externally validated soft causal priors than as direct evidence of causal structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。