解决因果推断中分组不一致导致的预测偏差问题。
Consistent Labeling Across Group Assignments: Variance Reduction in Conditional Average Treatment Effect Estimation
- 提出跨分组标签一致性机制,消除同一实例在不同分组下的预测矛盾。
- 实验表明,该方法在合成与真实数据集上均显著降低测试误差。
- 适用于所有已有CATE估计算法,特别适合需要高稳定性的医疗决策场景。
众多条件平均处理效应(CATE)估计算法存在一个共性问题:同一实例在不同分组分配下表现出不一致的学习行为。本文首次提出量化并可视化这种不一致性的度量方法,并通过理论分析证明该不一致性会带来更高的测试误差,且无法通过传统机器学习手段解决。为此,我们提出一种通用方法——跨分组标签一致性(CLAGA),可消除此类不一致,兼容任何现有CATE算法。在合成数据和真实世界数据集上的实验表明,引入CLAGA后模型性能显著提升。
原文摘要 · Abstract (English)
Numerous algorithms have been developed for Conditional Average Treatment Effect (CATE) estimation. In this paper, we first highlight a common issue where many algorithms exhibit inconsistent learning behavior for the same instance across different group assignments. We introduce a metric to quantify and visualize this inconsistency. Next, we present a theoretical analysis showing that this inconsistency indeed contributes to higher test errors and cannot be resolved through conventional machine learning techniques. To address this problem, we propose a general method called \textbf{Consistent Labeling Across Group Assignments} (CLAGA), which eliminates the inconsistency and is applicable to any existing CATE estimation algorithm. Experiments on both synthetic and real-world datasets demonstrate significant performance improvements with CLAGA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。