arXiv:2505.17133stat.MLcs.AI2025-05被引 6

用少量可靠数据训练模型,预测各类人群的因果概率边界。

Learning Probabilities of Causation with Mask-Augmented Data

  • 基于掩码增强数据训练神经网络,从局部数据推断全局因果概率
  • 在10万至20万样本数据上平均绝对误差仅0.03,较基线降低80%
  • 适合需要高精度因果推断的医疗、政策评估等决策场景

因果概率在现代决策中至关重要。Tian和Pearl首次给出了三种二值因果概率(如必要性与充分性概率PNS)的形式化定义及紧致边界。然而,估计这些概率需针对每个亚群体获取实验与观测分布,这在有限的群体层面数据下往往不可靠或不切实际。为此,我们提出两种机器学习模型:Exact-MLP与Mask-MLP,它们在少量可靠亚群体上训练,即可预测所有其他亚群体的PNS边界。我们在四个结构因果模型(SCMs)上验证了该方法,每项评估使用10万至20万样本的群体级数据。模型在主任务上平均绝对误差(MAE)约为0.03,相较对应基线减少约80%。结果表明,机器学习可有效学习因果概率,且所提方法具有可行性与高效性。

原文摘要 · Abstract (English)

Probabilities of causation play a central role in modern decision making. Tian and Pearl first introduced formal definitions and derived tight bounds for three binary probabilities of causation, such as the probability of necessity and sufficiency (PNS). However, estimating these probabilities requires both experimental and observational distributions specific to each subpopulation, which are often unreliable or impractical to obtain from limited population-level data. To solve this problem, we propose two machine learning models: Exact-MLP and Mask-MLP, which are trained on a small set of reliable subpopulations and are able to predict PNS bounds for all other subpopulations. We validate our models across four Structural Causal Models (SCMs), each evaluated on population-level data with sample sizes between 100k and 200k. Our models achieve average mean absolute errors (MAEs) of roughly 0.03 on main tasks, reducing MAE by about 80% relative to the corresponding baselines. These results demonstrate both the feasibility of machine learning models for learning probabilities of causation and the effectiveness of the proposed approach.

因果推断机器学习概率边界数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。