新方法解决弱监督学习中的类别先验估计难题。
Mixture Proportion Estimation and Weakly-supervised Kernel Test for Conditional Independence
- 基于条件独立性假设,突破传统不可约性限制。
- 提出方法矩估计器,理论分析其渐近性质。
- 配套核检验可验证假设,适合因果推断与公平性研究。
混合比例估计(MPE)旨在从无标签数据中估计类别先验,是弱监督学习中关键步骤,如PU学习、带标签噪声学习和域自适应。现有方法依赖不可约性假设或其变体以保证可识别性。本文提出基于给定类别标签的条件独立性假设,在不可约性不成立时仍能保证可识别性。在此基础上,我们构建了方法矩估计器,并分析其渐近性质。此外,还提出了弱监督核检验来验证条件独立性假设,该检验在因果发现与公平性评估等应用中具有独立价值。实验表明,所提估计器性能优于现有方法,且检验能有效控制第一类与第二类错误。
原文摘要 · Abstract (English)
Mixture proportion estimation (MPE) aims to estimate class priors from unlabeled data. This task is a critical component in weakly supervised learning, such as PU learning, learning with label noise, and domain adaptation. Existing MPE methods rely on the \textit{irreducibility} assumption or its variant for identifiability. In this paper, we propose novel assumptions based on conditional independence (CI) given the class label, which ensure identifiability even when irreducibility does not hold. We develop method of moments estimators under these assumptions and analyze their asymptotic properties. Furthermore, we present weakly-supervised kernel tests to validate the CI assumptions, which are of independent interest in applications such as causal discovery and fairness evaluation. Empirically, we demonstrate the improved performance of our estimators compared with existing methods and that our tests successfully control both type I and type II errors.\label{key}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。