利用相似性置信与置信差异双信号,提升弱监督学习性能
Learning from Similarity-Confidence and Confidence-Difference
- 融合相似性置信与置信差异两种弱标签信号
- 在多个数据集上显著优于现有基线方法
- 适合标注稀缺场景下的模型训练
在实际机器学习应用中,数据标注往往不准确,且标注样本数量有限。弱监督学习(WSL)通过利用不完整或不精确的监督信号提供有效解决方案。然而,现有方法多依赖单一类型弱监督信号。本文提出一种新框架,从多个关系视角整合互补的弱监督信号,尤其适用于标注数据稀缺的情况。我们设计SconfConfDiff分类方法,对无标签数据对赋予相似性置信和置信差异两类弱标签。为此,推导出两种无偏风险估计器:一种基于现有估计器的凸组合,另一种通过建模两弱标签交互关系新设计。理论证明二者均达到最优收敛速率。此外,提出风险校正策略以缓解负经验风险导致的过拟合,并分析了方法对错误类别先验和标签噪声的鲁棒性。实验表明,该方法在多种设置下持续优于现有基线。
原文摘要 · Abstract (English)
In practical machine learning applications, it is often challenging to assign accurate labels to data, and increasing the number of labeled instances is often limited. In such cases, Weakly Supervised Learning (WSL), which enables training with incomplete or imprecise supervision, provides a practical and effective solution. However, most existing WSL methods focus on leveraging a single type of weak supervision. In this paper, we propose a novel WSL framework that leverages complementary weak supervision signals from multiple relational perspectives, which can be especially valuable when labeled data is limited. Specifically, we introduce SconfConfDiff Classification, a method that integrates two distinct forms of weaklabels: similarity-confidence and confidence-difference, which are assigned to unlabeled data pairs. To implement this method, we derive two types of unbiased risk estimators for classification: one based on a convex combination of existing estimators, and another newly designed by modeling the interaction between two weak labels. We prove that both estimators achieve optimal convergence rates with respect to estimation error bounds. Furthermore, we introduce a risk correction approach to mitigate overfitting caused by negative empirical risk, and provide theoretical analysis on the robustness of the proposed method against inaccurate class prior probability and label noise. Experimental results demonstrate that the proposed method consistently outperforms existing baselines across a variety of settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。