arXiv:2605.08197cs.LGcs.AI2026-05

用可执行回放测试大模型从干预数据中推断因果机制的能力

ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions

  • 构建1300个二值世界,要求模型输出可验证的因果机制代码
  • 隐藏结构顺序或根节点时,模型泛化能力显著下降
  • 引入审计链提升推理可靠性,防止错误替代方案通过

现有因果推理基准多评估局部答案或图结构。我们提出ReplaySCM,一个包含1300个条目的基准,用于从有限干预证据中推断可执行的因果机制。每个条目由一个潜在的完全可观测的无环布尔结构性因果模型(SCM)生成。系统需输出受限布尔领域特定语言(DSL)的机制图;提交内容经解析、合法性与无环性检查后,在训练和保留的干预世界中回放。评分基于回放行为而非公式字符串,因此语法不同但行为正确的机制可获认可。ReplaySCM通过有序、块序、隐藏序和隐藏根四种设置改变模型可见的结构信息,并包含替代SCM任务:提供一个有效参考SCM,要求找出语义不同的替代方案及其分离干预与见证。前沿大模型能推断部分功能-父节点结构,但当顺序或根结构被隐藏时,保留回放性能显著下降。我们还评估了匹配的支持-审计阶梯:原始、额外世界、反例审计(CEx),将局部前驱模式覆盖率从0.8949提升至0.9815再至1.0;在审计搜索下,所有发现的语义替代方案均不再与训练世界一致。即使在更强证据下,有序/隐藏序差距依然存在。ReplaySCM通过评估从有限干预证据中的可执行回放泛化能力,补充了答案级因果推理与图发现基准,不声称唯一识别潜在SCM。

原文摘要 · Abstract (English)

Most causal benchmarks for language models score local answers or graph structure. We introduce ReplaySCM, a 1,300 item benchmark for executable causal mechanism induction from finite interventional evidence. Each item contains binary worlds generated by a latent fully observed acyclic Boolean structural causal model (SCM). A system must output a mechanism map in a restricted Boolean DSL; the submission is parsed, checked for legality and acyclicity, and replayed on training and held-out intervention worlds. Scoring uses replay behavior rather than formula strings, so syntactically different mechanisms receive credit when they behave correctly. ReplaySCM varies the structural information disclosed to the model through Ordered, Block-order, Hidden-order, and Hidden-roots settings, and includes Alternative-SCM tasks that supply a valid reference SCM and ask for a semantically distinct alternative that fits the training worlds, together with a separating intervention and witness. Frontier LLMs infer parts of the functional-parent structure, but held-out replay drops sharply when order or root structure is hidden. We also evaluate a matched support-audit ladder: Original, Extra Worlds, and Counterexample Audit (CEx), that raises mean local predecessor-pattern coverage from 0.8949 to 0.9815 to 1.0; under the audited searches, no discovered semantic alternative remains consistent with the training worlds. The Ordered/Hidden-order gap persists under this stronger evidence. ReplaySCM complements answer-level causal reasoning and graph-discovery benchmarks by evaluating executable replay generalization from finite interventional evidence, without claiming unique identification of the latent SCM.

因果推理大模型评估可执行回放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。