测试大模型在游戏中的因果推理能力,发现当前模型普遍表现不佳。
CausalGame: Benchmarking Causal Thinking of LLM Agents in Games

- 设计14个含选择偏差等真实科学挑战的游戏场景
- 30个模型最佳仅68%生存率,远低于最优解78-85%
- 适合评估大模型的科学推理能力,尤其关注因果思维
构建具备大型语言模型(LLM)的AI科学家代理近年来受到广泛关注。由于科学发现本质上依赖于从观察中揭示因果关系,因此因果思维能力——即区分因果与相关性、识别隐藏偏差——对LLM代理至关重要。尽管已有多个针对AI科学家的基准测试,但均未明确包含现实科学发现中广泛存在的选择偏差、测量误差和隐含混杂因素。为此,我们提出CausalGame,一个通过交互式游戏评估LLM代理因果思维能力的基准。CausalGame要求LLM代理主动设计实验方案、收集观测数据并生成最终解决方案及解释报告。为模拟真实的科学发现挑战,我们设计了14种包含选择偏差、测量误差和隐含混杂因素的场景。在30个LLM代理中,没有一个展现出可靠的因果思维:表现最好的模型在对抗分析最优解时仅达到68.0%的生存率,而78%-85%为理论最优水平;仅有5%-7%的会话在因果推理评分上获得认可。CausalGame为评估AI科学家代理的因果思维提供了一个可扩展且可控的测试平台。
原文摘要 · Abstract (English)
Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention. Since scientific discovery fundamentally relies on uncovering causal relationships from observations, the capability of causal thinking, i.e., distinguishing causation from correlation and recognizing hidden biases, is essential to LLM agents. Although a number of benchmarks exist for AI Scientists, none explicitly incorporate challenges from selection bias, measurement error, and hidden confounders that widely exist in real-world scientific discovery. To this end, we present CausalGame, a benchmark that evaluates the causal thinking capabilities of LLM agents through interactive games. CausalGame asks LLM agents to actively design experimental protocols, collect observation data, and derive a final solution with an explanation report. To emulate realistic scientific discovery challenges, we design 14 scenarios that incorporate selection bias, measurement error, and hidden confounders. Across 30 LLM agents, none demonstrates reliable causal thinking: the best model reaches only 68.0% survival against analytical optima of 78-85%, and merely 5-7% of sessions receive credits on the causal-reasoning rubrics. CausalGame provides a scalable and controlled testbed for evaluating the causal thinking of AI Scientist agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。