arXiv:2601.21210cs.AI2026-01Conference of the …被引 3

用符号验证找回大模型因果推理中被误判的正确答案

Uncovering Hidden Correctness in LLM Causal Reasoning via Symbolic Verification

  • 通过因果图与do演算规则验证模型输出的逻辑可推导性
  • 在合成数据和问答基准上显著提升语义正确性的识别率
  • 适合关注模型推理可信度的开发者与研究者使用

大型语言模型(LLMs)在因果推理任务中的应用日益广泛。然而,现有评估基准多依赖字符串匹配或表层指标,无法判断模型输出在因果语义上的形式有效性。为此,我们提出DoVerifier,一种基于符号的验证器,通过因果图与do演算、概率论规则,检验LLM生成的因果表达是否可推导。该方法能恢复原本因表面差异而被判定为错误的正确答案。在合成数据和因果问答基准上的评估表明,DoVerifier更准确地捕捉了因果推理轨迹的语义正确性,为评估LLMs在因果推理任务中的表现提供了更严格、更富有信息量的方法。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being applied to tasks that involve causal reasoning. However, current benchmarks often rely on string matching or surface-level metrics that do not capture whether the output of a model is formally valid under the semantics of causal reasoning. To address this, we propose DoVerifier, a simple symbolic verifier that checks whether LLM-generated causal expressions are derivable from a given causal graph using rules from do-calculus and probability theory. This allows us to recover correct answers to causal queries that would otherwise be marked incorrect due to superficial differences in their causal semantics. Our evaluations on synthetic data and causal QA benchmarks show that DoVerifier more accurately captures semantic correctness of causal reasoning traces, offering a more rigorous and informative way to evaluate LLMs on causal reasoning.

因果推理符号验证大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。