arXiv:2411.11256cs.LGstat.ML2024-11KDD被引 4

小样本下主动学习提升因果效应估计的泛化能力

Progressive Generalization Risk Reduction for Data-Efficient Causal Effect Estimation

  • 逐步缩小泛化风险上界,设计可解释的标签获取策略
  • 在有限标注预算下,显著降低因果效应估计误差
  • 适合医疗等数据稀缺但高风险场景的决策支持系统

因果效应估计(CEE)是预测实体未观测反事实结果的关键工具。由于无需理想化的完全匹配样本,而可基于观察数据进行推断,其广泛应用于医疗治疗效果预测等高风险领域。然而,在这些领域中,获取足够规模且完全标注的观察数据仍面临成本、伦理、专业性和时间等多重挑战,例如医学调查。因此,当训练数据量较小时,现有CEE算法难以实现低泛化风险。与以往假设数据充足不同,本文研究更现实的设定:初始标注样本稀少,但可在训练过程中逐步获取更多数据(受限于预算)。核心问题在于如何主动选择最优样本进行标注,例如确定下一步应调查哪些患者。然而,在有限标注预算下,有效降低CEE风险的方法尚未充分探索。为此,本文从渐进式缩小泛化风险上界的视角进行理论分析,提出专为CEE任务设计的原理性标签获取流程。基于该分析,我们提出模型无关的因果主动学习(MACAL)算法,以批处理方式选取样本,旨在同时减少模型不确定性与标注后风险。

原文摘要 · Abstract (English)

Causal effect estimation (CEE) provides a crucial tool for predicting the unobserved counterfactual outcome for an entity. As CEE relaxes the requirement for ``perfect'' counterfactual samples (e.g., patients with identical attributes and only differ in treatments received) that are impractical to obtain and can instead operate on observational data, it is usually used in high-stake domains like medical treatment effect prediction. Nevertheless, in those high-stake domains, gathering a decently sized, fully labelled observational dataset remains challenging due to hurdles associated with costs, ethics, expertise and time needed, etc., of which medical treatment surveys are a typical example. Consequently, if the training dataset is small in scale, low generalization risks can hardly be achieved on any CEE algorithms. Unlike existing CEE methods that assume the constant availability of a dataset with abundant samples, in this paper, we study a more realistic CEE setting where the labelled data samples are scarce at the beginning, while more can be gradually acquired over the course of training -- assuredly under a limited budget considering their expensive nature. Then, the problem naturally comes down to actively selecting the best possible samples to be labelled, e.g., identifying the next subset of patients to conduct the treatment survey. However, acquiring quality data for reducing the CEE risk under limited labelling budgets remains under-explored until now. To fill the gap, we theoretically analyse the generalization risk from an intriguing perspective of progressively shrinking its upper bound, and develop a principled label acquisition pipeline exclusively for CEE tasks. With our analysis, we propose the Model Agnostic Causal Active Learning (MACAL) algorithm for batch-wise label acquisition, which aims to reduce both the CEE model's uncertainty and the post-acquisition ...

因果推断主动学习小样本医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。