arXiv:2509.04112cs.LGcs.IT2025-09被引 1

用合成数据提升反事实预测区间的精度与效率

Synthetic Counterfactual Labels for Efficient Conformal Counterfactual Inference

  • 通过预训练模型生成合成反事实标签,扩充校准集
  • 在多种设置下显著缩小预测区间,同时保持覆盖率
  • 适合需要精准个体化反事实推断的研究者

本文针对个体反事实结果的可靠预测区间构建问题,提出合成数据驱动的反事实推断框架(SP-CCI)。现有方法虽能提供边际覆盖率保证,但在处理治疗不平衡时因反事实样本稀少而产生过保守的区间。SP-CCI利用预训练反事实模型生成合成标签,扩充校准集,并结合风险控制预测集(RCPS)与预测驱动推断(PPI)的去偏步骤,确保方法有效性。理论证明其在精确与近似重要性加权下均能获得更紧的预测区间且保持边际覆盖率。实验证明,在多个数据集上,SP-CCI始终比标准方法更小的区间宽度。

原文摘要 · Abstract (English)

This work addresses the problem of constructing reliable prediction intervals for individual counterfactual outcomes. Existing conformal counterfactual inference (CCI) methods provide marginal coverage guarantees but often produce overly conservative intervals, particularly under treatment imbalance when counterfactual samples are scarce. We introduce synthetic data-powered CCI (SP-CCI), a new framework that augments the calibration set with synthetic counterfactual labels generated by a pre-trained counterfactual model. To ensure validity, SP-CCI incorporates synthetic samples into a conformal calibration procedure based on risk-controlling prediction sets (RCPS) with a debiasing step informed by prediction-powered inference (PPI). We prove that SP-CCI achieves tighter prediction intervals while preserving marginal coverage, with theoretical guarantees under both exact and approximate importance weighting. Empirical results on different datasets confirm that SP-CCI consistently reduces interval width compared to standard CCI across all settings.

反事实推断预测区间合成数据机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。