arXiv:2507.14528cs.LG2025-07中稿 · KDD

用已知治疗组识别控制组,解决观察性研究中无明确对照的难题。

Positive-Unlabeled Learning for Control Group Construction in Observational Causal Inference

  • 仅凭已知治疗样本,通过正负样本学习筛选高置信度控制样本。
  • 在模拟与真实农业数据中,重构的对照组使因果效应估计接近真实值。
  • 适用于难以随机试验的环境、农业等领域,尤其适合有观测数据但缺对照的情况。

在因果推断中,无论是随机试验还是观察性研究,获取治疗组和对照组都至关重要。随机分配下可直接比较组间结果估算平均处理效应(ATE)。非随机情形需调整混杂因素以逼近反事实情景。观察性研究常见难题是缺乏明确标记为未接受治疗的对照单位。为此,我们提出正负样本(PU)学习框架,仅利用已知治疗单位(正样本),从无标签样本中高置信度识别控制单位。我们在模拟与真实世界数据上评估该方法:构建包含多样关系的因果图生成合成数据,在不同场景下检验其恢复对照组并准确估计真实 ATE 的能力。还将方法应用于可持续农业中的最优播种与施肥方案真实数据。结果表明,该方法能仅基于治疗单位成功识别控制单位,并由此估计出接近真实值的 ATE。这项工作对观察性因果推断具有重要意义,尤其在随机实验困难或成本高的领域。在地球、环境与农业科学中,可通过利用现有遥感与气候数据,实现大量准实验研究,尤其当治疗单位存在而对照单位缺失时。

原文摘要 · Abstract (English)

In causal inference, whether through randomized controlled trials or observational studies, access to both treated and control units is essential for estimating the effect of a treatment on an outcome of interest. When treatment assignment is random, the average treatment effect (ATE) can be estimated directly by comparing outcomes between groups. In non-randomized settings, various techniques are employed to adjust for confounding and approximate the counterfactual scenario to recover an unbiased ATE. A common challenge, especially in observational studies, is the absence of units clearly labeled as controls-that is, units known not to have received the treatment. To address this, we propose positive-unlabeled (PU) learning as a framework for identifying, with high confidence, control units from a pool of unlabeled ones, using only the available treated (positive) units. We evaluate this approach using both simulated and real-world data. We construct a causal graph with diverse relationships and use it to generate synthetic data under various scenarios, assessing how reliably the method recovers control groups that allow estimates of true ATE. We also apply our approach to real-world data on optimal sowing and fertilizer treatments in sustainable agriculture. Our findings show that PU learning can successfully identify control (negative) units from unlabeled data based only on treated units and, through the resulting control group, estimate an ATE that closely approximates the true value. This work has important implications for observational causal inference, especially in fields where randomized experiments are difficult or costly. In domains such as earth, environmental, and agricultural sciences, it enables a plethora of quasi-experiments by leveraging available earth observation and climate data, particularly when treated units are available but control units are lacking.

因果推断正负学习农业科学观察研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。