在干预实验中,通过部分因果结构学习,提升选择性置信预测的准确性。
Partial Causal Structure Learning for Valid Selective Conformal Inference under Interventions
- 只学习干预与目标变量的关系,而非完整因果图
- 估计错误会导致覆盖率下降,可提供保守修正方案
- 适用于基因组扰动实验等真实干预场景
选择性置信预测可在识别出与测试样本可交换的校准样本时,显著缩小不确定性区间。在干预场景(如基因组扰动实验)中,这种可交换性通常仅在不影响目标变量的干预子集中成立(例如因果图中未受干预节点影响的非后代)。我们研究实际情形下该不变性结构未知、需从数据中估计的问题。主要结果量化了当估计的校准集意外包含影响目标变量的干预时,覆盖率下降的程度,并在已知误差上限时提供保守修正方法。我们不学习完整因果图,仅学习决定校准干预所需的干预-目标关系。提出了相应的算法,并在合成结构方程模型和Replogle K562 CRISPR干扰数据上进行评估,实验展示了选择性校准带来的合成增益及真实扰动筛选中的有限样本权衡。
原文摘要 · Abstract (English)
Selective conformal prediction can yield substantially tighter uncertainty sets when we can identify calibration examples that are exchangeable with the test example. In interventional settings, such as perturbation experiments in genomics, exchangeability often holds only within subsets of interventions that leave a target variable "unaffected" (e.g., non-descendants of an intervened node in a causal graph). We study the practical regime where this invariance structure is unknown and must be estimated from data. Our main result quantifies how coverage degrades when the estimated safe calibration set accidentally includes interventions that affect the target, and gives a conservative correction when an upper bound on this error is available. Rather than learning a full causal graph, we learn only the intervention-target relationships needed to choose calibration interventions. We give algorithms for this partial learning task and evaluate them on synthetic structural equation models and Replogle K562 CRISPR-interference data, where the experiments illustrate synthetic gains from selective calibration and finite-sample tradeoffs on real perturbation screens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。