利用干净样本提升部分标签学习的准确率
Exploiting the Potential Supervision Information of Clean Samples in Partial Label Learning
- 基于邻居样本关系,用干净样本校准标签置信度
- 在5个真实数据集上显著提升主流方法性能
- 适合处理标签不完整或含噪声的数据场景
降低误标标签的影响是解决部分标签学习中歧义问题的关键。现有方法多关注单个样本特征,忽视了数据集中随机存在的干净样本所蕴含的强监督信息。本文提出一种新校准策略 CleanSE,利用干净样本引导标签选择:假设每个干净样本若其标签属于其近邻样本的候选标签,则更可能是该近邻的真实标签。同时,通过限制每类标签数量在特定区间,帮助刻画样本分布。在3个合成基准和5个真实部分标签学习数据集上的大量实验表明,该策略可适配多数先进方法并有效提升性能。
原文摘要 · Abstract (English)
Diminishing the impact of false-positive labels is critical for conducting disambiguation in partial label learning. However, the existing disambiguation strategies mainly focus on exploiting the characteristics of individual partial label instances while neglecting the strong supervision information of clean samples randomly lying in the datasets. In this work, we show that clean samples can be collected to offer guidance and enhance the confidence of the most possible candidates. Motivated by the manner of the differentiable count loss strat- egy and the K-Nearest-Neighbor algorithm, we proposed a new calibration strategy called CleanSE. Specifically, we attribute the most reliable candidates with higher significance under the assumption that for each clean sample, if its label is one of the candidates of its nearest neighbor in the representation space, it is more likely to be the ground truth of its neighbor. Moreover, clean samples offer help in characterizing the sample distributions by restricting the label counts of each label to a specific interval. Extensive experiments on 3 synthetic benchmarks and 5 real-world PLL datasets showed this calibration strategy can be applied to most of the state-of-the-art PLL methods as well as enhance their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。