诊断大模型因果推理失败的5147个案例,揭示隐藏错误模式。
CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs
- 构建跨三层因果阶梯的诊断基准,标注失败原因与陷阱类型。
- 发现模型在压力下易产生错误判断,拒绝正确主张比例超30%。
- 适合研究可信因果推理的学者与开发安全模型的工程师使用。
大型语言模型日益生成流畅的因果解释,但其错误常无法被整体准确率捕捉:混淆关联与干预、在压力下放弃正确判断、过度拒绝有效主张或在证据不足时作答。我们提出CTK,一个包含5,147个案例并持续增长的诊断基准,覆盖10个领域及珍珠因果阶梯的全部三个层级。不同于仅评估正确性的基准,CTK通过标注因果层级、陷阱类型、压力敏感性、拒绝质量与效用-安全权衡,揭示模型失败原因。其羊/狼分类法区分有效因果设计与推断陷阱;成对的中性/压力变体测量‘谄媚漂移’(坏翻转率);智慧拒绝测试则检验模型识别缺失信息的能力。CTK暴露了被整体准确率掩盖的故障模式:怀疑陷阱、规模扩大下的层级坍塌、压力诱导漂移、检测-纠正差距及反事实错误模式。该基准不提供修正方法,而是为研究因果推理失败特征提供诊断基础。
原文摘要 · Abstract (English)
Large language models increasingly produce fluent causal explanations, yet they often fail in ways aggregate accuracy cannot diagnose: confusing association with intervention, abandoning correct judgments under pressure, over-refusing valid claims, or answering when evidence is underdetermined. We introduce CTK, a diagnostic benchmark of 5,147 cases and growing, across 10 domains and all three levels of Pearl's Ladder of Causation. Unlike benchmarks that only score correctness, CTK reveals why a model failed by annotating causal rung, trap type, pressure sensitivity, refusal quality, and Utility-Safety tradeoffs. Its Sheep/Wolf taxonomy separates valid causal designs from inferential traps; paired neutral/pressure variants measure sycophantic drift through Bad Flip Rate; and Wise Refusal fields test whether a model identifies the missing information needed before endorsing a claim. CTK exposes failure modes hidden by aggregate accuracy: the Skepticism Trap, Rung Collapse under scaling, pressure-induced drift, Detection-Correction gaps, and counterfactual error modes. Rather than prescribing a correction method, it provides the diagnostic substrate for studying causal-reasoning failure profiles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。