arXiv:2608.27953cs.AI2026-08中稿 · EMNLP

测试大模型对开放式反事实推理的能力,发现最强模型仅达64.62分。

The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs

论文配图:The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs
图 1 · 摘自论文原文
  • 构建220个跨领域反事实问题的诊断基准,支持长时序因果推理。
  • 提出PRISM评估框架,通过语义因果图量化解释的因果有效性。
  • 揭示模型常有流畅叙事但缺乏真实因果链条,适合研究因果推理的学者。

反事实推理要求模型超越观测世界,解释条件变化如何引发下游后果。现有基准多聚焦变量固定或单一答案的封闭场景,忽视需因果过程评估的开放域问题。为此,我们提出WhatIfBench,一个针对开放域、开放形式、长时程反事实因果推理的诊断基准,包含220个涵盖STEM、HSS和混合领域的what-if问题。为评估自由文本回答,我们进一步提出PRISM,首次将自然语言解释转化为响应衍生的语义因果图(事件、状态与机制),并在该图上联合应用过程度量(评估图级因果有效性)与评分度量(评估答案级解释充分性)。使用该框架评估六款前沿大模型,发现WhatIfBench尚未饱和:即使最强模型也仅达64.62%最终得分。深入分析显示存在持续的因果缺口、前提漂移与拓扑碎片化,表明流畅的反事实叙述常掩盖脆弱的因果过程。基准、代码与评估脚本已公开于WhatIfBench。

原文摘要 · Abstract (English)

Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks largely target bounded settings with fixed variables or single gold outcomes, overlooking open-domain scenarios requiring causal-process evaluation. To this end, we present $\textbf{WhatIfBench}$, a diagnostic benchmark for open-domain, open-form, long-horizon counterfactual causal reasoning, containing 220 what-if questions across STEM, HSS, and Hybrid scenarios. To evaluate free-form responses, we further propose $\textbf{PRISM}$, which first converts each natural-language explanation into a Response-Derived Semantic Causal Graph of events, states, and mechanisms. On top of this graph, PRISM then jointly applies a Process Metric assessing graph-level causal validity and a Rubric Metric assessing answer-level explanatory adequacy. Evaluating six frontier LLMs with this framework, we find that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score. Further analysis reveals persistent causal gaps, premise drift, and topology fragmentation, suggesting that fluent counterfactual narratives often mask fragile causal processes. The benchmark, code, and evaluation scripts are available at $\href{https://github.com/zju-gt/WhatIfBench}{WhatIfBench}$.

反事实推理因果建模大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。