构建首个融合因果推理与数据科学任务的综合评测基准。
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

- 基于生成式结构因果模型和自然语言故事构建真实场景
- 涵盖珍珠因果三阶梯,含编码与不确定度评估任务
- 适合评估智能体在复杂数据任务中的理性决策能力
大型语言模型正日益作为集成的数据科学智能体,结合抽象推理与工具使用。然而现有评测体系要么仅关注符号化因果推理而缺乏真实数据分析,要么仅有数据处理任务但缺少严谨的因果生成结构。现有因果数据集多来自有限模板化的已有样本,缺乏系统性合成的新因果结构。本文提出CausalDS,一个用于评估智能体在数据科学工作流中因果推理能力的基准。每个测试实例包含一个采样的结构因果模型(SCM)及其生成的观测数据,以及一个基于真实领域的合成自然语言故事。可选地,将基准组件的组合基于真实数据集的分布,保留实证结构的同时通过完全合成避免“因果鹦鹉”问题。从每个场景中衍生出涵盖珍珠因果三阶梯的任务,典型预测任务为第一阶梯。多数任务包含数据科学编程环节,模型需调用多个工具以应对不完整观测(由观测模型生成)。同时,识别无法合理回答的问题并选择放弃,也被视为首类评分结果。该基准共同评估符号因果推理、数据科学、不确定性量化、拒绝回答及工具使用/编程能力。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure. Furthermore, existing causal evaluation datasets are often restricted to curated examples from existing sources, with diversity coming from limited templatized variations rather than from systematic generation of novel synthetic causal structures. We introduce CausalDS, a benchmark for evaluating causal reasoning in agentic data-science workflows. Each benchmark instance is a scene consisting of a sampled structural causal model (SCM) with generated observational data and an accompanying synthetic natural-language story grounded in a realistic domain. We optionally ground the composition of the benchmark components in empirical distributions obtained from real-world datasets, thus retaining empirical structure while reducing the "causal parrot" risk through completely synthetic generation. From each scene, we then derive tasks spanning all three of Pearl's rungs, with typical data-science prediction tasks appearing as Rung 1. Most tasks include a data science coding component, where the model typically needs to use several tools to arrive at the final answer due to the frequent presence of imperfect observations, which are generated by an observation model. Additionally, recognizing when a question admits no warranted answer and abstaining is treated as a first-class scored outcome. The benchmark thus jointly evaluates symbolic causal reasoning, data science, uncertainty quantification, abstention, and tool use/coding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。