arXiv:2411.03021cs.LGstat.AP2024-11

提出可统计评估因果模型泛化能力的新框架,解决真实场景下模型可靠性验证难题。

Testing Generalizability in Causal Inference

  • 基于简约参数化构建全/半合成因果基准,灵活模拟跨域数据偏移
  • 在真实世界数据基础上评估,避免传统方法依赖简化数据集的偏差
  • 结合模拟与统计检验,为模型选择提供可信赖的决策依据

确保机器学习模型在多样化现实场景中表现稳健,需解决因协变量偏移导致的泛化性问题。然而,当前尚无针对机器学习算法泛化性的正式统计评估流程。现有预测指标如均方误差(MSE)仅能比较模型相对性能,无法直接判断模型是否具备泛化能力。针对因果推断领域这一空白,本文提出一种系统性框架,用于统计评估高维因果推断模型的泛化能力。该方法采用简约参数化,灵活生成全合成与半合成因果基准,实现对均值与分布回归方法的全面评估。其基于真实世界数据,提升了评估的真实性,弥补了当前研究多依赖简化数据集的不足。此外,通过模拟与统计检验相结合,框架具备鲁棒性,减少对传统指标的过度依赖,为决策提供统计保障。

原文摘要 · Abstract (English)

Ensuring robust model performance in diverse real-world scenarios requires addressing generalizability across domains with covariate shifts. However, no formal procedure exists for statistically evaluating generalizability in machine learning algorithms. Existing predictive metrics like mean squared error (MSE) help to quantify the relative performance between models, but do not directly answer whether a model can or cannot generalize. To address this gap in the domain of causal inference, we propose a systematic framework for statistically evaluating the generalizability of high-dimensional causal inference models. Our approach uses the frugal parameterization to flexibly simulate from fully and semi-synthetic causal benchmarks, offering a comprehensive evaluation for both mean and distributional regression methods. Grounded in real-world data, our method ensures more realistic evaluations, which is often missing in current work relying on simplified datasets. Furthermore, using simulations and statistical testing, our framework is robust and avoids over-reliance on conventional metrics, providing statistical safeguards for decision making.

因果推断泛化性评估统计检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。