测试大模型在真实因果实验中的理解能力,发现能猜对结果却不明白原理。
CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

- 用模拟实验室让大模型做因果推断实验,需同时预测和还原机制
- 模型预测准确率达92%,但因果图正确率仅0.471,差距明显
- 适合研究大模型因果推理能力的学者,尤其关注可解释性者
我们提出CausaLab,一个用于评估大模型代理进行交互式因果发现的可扩展环境。与以往评估不同,CausaLab不仅检验代理能否利用因果证据解决问题,还考察其答案是否基于真实恢复的因果机制。每个实验将代理置于合成实验室中:接收历史测量数据,干预操控晶体,预测未观测反应晶体的共振频率,该频率由相同机制决定。隐藏的数据生成过程为随机采样的结构因果模型(SCM),因此成功要求同时恢复因果图和结构方程,而非依赖已有知识。实验显示,预测与机制恢复之间存在持续差距:在纯观测6节点设置下,GPT-5.2-high任务准确率达92%,但所有边的F₁仅为0.471。混合观测-干预策略提升结构保真度,而纯干预仍对强模型构成挑战。我们识别出过早停止是主要弱点,并证明一致性验证可缓解此问题。CausaLab因此区分了预测成功与因果理解,揭示了当前大模型作为实验性因果推理者的局限。
原文摘要 · Abstract (English)
We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates both whether an agent can solve a problem using causal evidence and whether its answer is grounded in a faithful recovered causal mechanism. Each episode places an agent in a synthetic laboratory: it receives prior measurement records, intervenes on a manipulator crystal, and predicts the resonance frequency of a held-out reactor crystal governed by the same mechanism. The hidden data-generating process is a randomly sampled structural causal model (SCM), so success requires recovering both a causal graph and structural equations rather than recalling prior knowledge. Experiments show a persistent gap between prediction and mechanism recovery: in the purely observational 6-node setting, GPT-5.2-high reaches 92% task accuracy but only 0.471 all-edge $F_1$. Mixed observation-intervention strategies improve structural fidelity, while pure intervention remains difficult even for strong agents. We identify premature stopping as a major weakness and show that consistency verification mitigates it. CausaLab therefore separates predictive success from causal understanding and exposes current LLM agents' limits as experimental causal reasoners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。