提出MDPU框架,用不等式约束提升弱监督分类效果
Learning from M-Tuple Dominant Positive and Unlabeled Data
- 基于正例不少于负例的约束建模任意大小元组分布
- 推导出无偏风险估计器并证明其泛化误差界
- 适合标注比例不精确的现实场景,如医疗数据
标签比例学习(LLP)解决多实例分组为包的分类问题,每个包仅提供各类别实例的比例信息。但在实际应用中,精确获取特定类别实例比例非常困难。为此,本文提出一种广义学习框架MDPU,首先在正例数量不低于负例数量的约束下,对任意大小元组内实例分布进行数学建模;随后基于经验风险最小化(ERM)方法,推导出满足风险一致性的无偏风险估计器;为缓解训练过程中的过拟合问题,引入风险校正方法,得到修正后的风险估计器。理论分析表明,该无偏风险估计器具有良好的泛化误差界。在多个数据集上的大量实验及与现有基线方法的对比,全面验证了所提框架的有效性。
原文摘要 · Abstract (English)
Label Proportion Learning (LLP) addresses the classification problem where multiple instances are grouped into bags and each bag contains information about the proportion of each class. However, in practical applications, obtaining precise supervisory information regarding the proportion of instances in a specific class is challenging. To better align with real-world application scenarios and effectively leverage the proportional constraints of instances within tuples, this paper proposes a generalized learning framework \emph{MDPU}. Specifically, we first mathematically model the distribution of instances within tuples of arbitrary size, under the constraint that the number of positive instances is no less than that of negative instances. Then we derive an unbiased risk estimator that satisfies risk consistency based on the empirical risk minimization (ERM) method. To mitigate the inevitable overfitting issue during training, a risk correction method is introduced, leading to the development of a corrected risk estimator. The generalization error bounds of the unbiased risk estimator theoretically demonstrate the consistency of the proposed method. Extensive experiments on multiple datasets and comparisons with other relevant baseline methods comprehensively validate the effectiveness of the proposed learning framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。