arXiv:2511.22842cs.LGcs.AI2025-11中稿 · ICML

用随机生成的合成数据评估因果机器学习模型,提升实验透明度与可靠性。

CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning

  • 基于显式设计原则生成包含因果模型、数据和查询的合成基准
  • 可在观测、干预、反事实三类因果推理下进行严格评估
  • 适合研究者验证算法在不同假设下的鲁棒性

因果机器学习旨在通过机器学习算法回答“如果……会怎样”的问题,是高风险决策的重要工具。然而当前因果机器学习的实证评估方法仍十分有限,现有基准多依赖少量手工或半合成数据集,导致结论脆弱且不可推广。为此,我们提出CausalProfiler——一个用于因果机器学习方法的合成基准生成器。该工具基于对因果模型、查询及数据类别的明确设计选择,随机采样因果模型、数据、查询与真实答案,构建合成因果基准。通过这种方式,可对因果机器学习方法在多种条件下进行严谨且透明的评估。本工作首次实现具备覆盖保证与透明假设的合成因果基准随机生成,涵盖观测、干预与反事实三个层次的因果推理。我们通过在多种条件与假设下评估多个前沿方法,展示了CausalProfiler所能支持的分析类型与洞见。

原文摘要 · Abstract (English)

Causal machine learning (Causal ML) aims to answer "what if" questions using machine learning algorithms, making it a promising tool for high-stakes decision-making. Yet, empirical evaluation practices in Causal ML remain limited. Existing benchmarks often rely on a handful of hand-crafted or semi-synthetic datasets, leading to brittle, non-generalizable conclusions. To bridge this gap, we introduce CausalProfiler, a synthetic benchmark generator for Causal ML methods. Based on a set of explicit design choices about the class of causal models, queries, and data considered, the CausalProfiler randomly samples causal models, data, queries, and ground truths constituting the synthetic causal benchmarks. In this way, Causal ML methods can be rigorously and transparently evaluated under a variety of conditions. This work offers the first random generator of synthetic causal benchmarks with coverage guarantees and transparent assumptions operating on the three levels of causal reasoning: observation, intervention, and counterfactual. We demonstrate its utility by evaluating several state-of-the-art methods under diverse conditions and assumptions, both in and out of the identification regime, illustrating the types of analyses and insights the CausalProfiler enables.

因果推理合成数据基准评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。