arXiv:2505.05242cs.LG2025-05ICML被引 5

用反事实覆盖思想提升少标签下的治疗效果估计精度

Enhancing Treatment Effect Estimation via Active Learning: A Counterfactual Covering Perspective

  • 基于反事实覆盖半径优化,主动选择最有价值的标注样本
  • 在合成与半合成数据上显著优于现有基线方法
  • 适合医疗等标注成本高的场景,尤其肿瘤疗效评估

尽管近年来已发展出众多复杂的治疗效果估计算法,但在标注数据不足时性能仍受限,因治疗后效果标注成本高昂(如昂贵的肿瘤影像或活检)。因此,必须在有限标注预算下高效引入高质量标签。本文从主动学习视角出发,通过理论分析提出关键指标——事实与反事实覆盖半径,决定风险上界。为缩小该上界,提出贪心半径缩减算法,在理想平衡数据分布下表现优异;为进一步适应现实数据分布,提出FCCM方法,将优化目标转化为事实与反事实覆盖最大化,确保数据获取过程中的半径有效缩减。在全合成与半合成数据集上的基准测试表明,FCCM优于多种基线方法。

原文摘要 · Abstract (English)

Although numerous complex algorithms for treatment effect estimation have been developed in recent years, their effectiveness remains limited when handling insufficiently labeled training sets due to the high cost of labeling the effect after treatment, e.g., expensive tumor imaging or biopsy procedures needed to evaluate treatment effects. Therefore, it becomes essential to actively incorporate more high-quality labeled data, all while adhering to a constrained labeling budget. To enable data-efficient treatment effect estimation, we formalize the problem through rigorous theoretical analysis within the active learning context, where the derived key measures -- \textit{factual} and \textit{counterfactual covering radius} determine the risk upper bound. To reduce the bound, we propose a greedy radius reduction algorithm, which excels under an idealized, balanced data distribution. To generalize to more realistic data distributions, we further propose FCCM, which transforms the optimization objective into the \textit{Factual} and \textit{Counterfactual Coverage Maximization} to ensure effective radius reduction during data acquisition. Furthermore, benchmarking FCCM against other baselines demonstrates its superiority across both fully synthetic and semi-synthetic datasets.

治疗效果估计主动学习反事实医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。