用不确定性推理生成更真实因果评估数据,提升估计器比较可靠性。
Improving Generative Methods for Causal Evaluation via Simulation-Based Inference
- 将生成方法与参数视为不确定,通过后验推断选择合适配置
- 生成数据的因果估计与真实数据高度一致,误差显著降低
- 适合需要可靠评估因果模型的研究者,尤其关注不确定性
生成能准确反映现实观测数据的合成数据集对于评估因果估计器至关重要,但仍是难题。现有生成方法虽能在观测数据基础上生成变化参数(如处理效应、混杂偏倚)的合成数据,但难以确定应使用何种方法及参数值,且通常需用户指定固定参数点估计,无法表达不确定性,限制了后验推断能力,可能导致评估不可靠。本文提出基于模拟的因果评估框架SBICE,将生成方法及其参数视为不确定,基于源数据推断其后验分布。借助模拟推断技术,SBICE识别合适的生成方法并推断参数分布,生成与源数据分布高度匹配的合成数据。实验表明,SBICE生成的数据能产生与源数据接近的因果估计,显著提升了估计器评估的可靠性,是一种鲁棒且考虑不确定性的因果评估方法。
原文摘要 · Abstract (English)
Generating synthetic datasets that accurately reflect real-world observational data is critical for evaluating causal estimators, but it remains a challenging task. Existing generative methods offer a solution by producing synthetic datasets anchored in the observed data (source data) while allowing variation in key parameters such as the treatment effect and amount of confounding bias. However, it is often unclear which generative methods to use and which values of parameters to choose when generating synthetic datasets. Moreover, existing methods typically require users to provide fixed point estimates of such parameters. This denies users the ability to express uncertainty over both generative methods and parameter values and removes the potential for posterior inference, potentially leading to unreliable estimator comparisons. We introduce simulation-based inference for causal evaluation (SBICE), a framework that treats the generative method and its corresponding generative parameters as uncertain and infers their posterior distribution given a source dataset. Leveraging techniques in simulation-based inference, SBICE identifies suitable generative methods and infers distributions over its parameter configurations to produce synthetic datasets closely aligned with the source data distribution. Empirical results demonstrate that SBICE improves the reliability of estimator evaluations by generating realistic datasets whose causal estimates closely match the estimates of the source data, making it a robust and uncertainty-aware approach to selecting causal estimators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。