arXiv:2503.04556cs.CLcs.AI2025-03ICML被引 6

评估大模型在因果与组合推理上的表现,发现复杂路径下错误率上升。

Compositional Causal Reasoning Evaluation in Language Models

  • 构建统一框架,同时衡量因果与组合推理能力
  • 在数学应用题中发现多种错误模式,复杂路径错误更明显
  • 适用于评估大模型因果推理能力,尤其关注路径复杂性影响

因果推理和组合推理是人工智能的两大核心目标。衡量这些行为的程度需要严谨的评估方法。我们提出一种统一视角,将两者结合,称为组合因果推理(CCR):即推断因果效应如何组合,或因果量如何在图结构中传播。我们构建了一个系统化评估框架,用于平均处理效应和必要充分概率的CCR评估。以证明为例,我们在LLama、Phi和GPT系列语言模型上展示了该框架的应用。在一道数学应用题中,框架揭示了多种分类不同的错误模式。所有模型在因果路径复杂度增加时,其CCR错误均上升,但o1模型例外。

原文摘要 · Abstract (English)

Causal reasoning and compositional reasoning are two core aspirations in AI. Measuring the extent of these behaviors requires principled evaluation methods. We explore a unified perspective that considers both behaviors simultaneously, termed compositional causal reasoning (CCR): the ability to infer how causal measures compose and, equivalently, how causal quantities propagate through graphs. We instantiate a framework for the systematic evaluation of CCR for the average treatment effect and the probability of necessity and sufficiency. As proof of concept, we demonstrate CCR evaluation for language models in the LLama, Phi, and GPT families. On a math word problem, our framework revealed a range of taxonomically distinct error patterns. CCR errors increased with the complexity of causal paths for all models except o1.

因果推理大模型评估组合推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。