arXiv:2506.19496cs.LG2025-06IJCAI

针对噪声标签导致的模型性能下降,提出可恢复与优化的训练框架

COLUR: Confidence-Oriented Learning, Unlearning and Relearning with Noisy-Label Data for Model Restoration and Refinement

  • 基于信心导向设计学习-遗忘-重学三阶段流程
  • 在四个真实数据集上均超越现有最优方法
  • 适合修复因标签错误导致性能下降的模型

大规模深度学习模型在多种任务中取得了显著成功。然而,若模型需在包含噪声标签(误导或模糊信息)的数据集上训练,其性能可能大幅下降。目前对模型因噪声标签导致性能退化后的恢复机制研究仍有限。受神经科学中‘遗忘机制’启发——通过遗忘错误知识来加速正确知识的重学,我们提出一种鲁棒的模型恢复与精炼(MRR)框架COLUR,即信心导向的学习、遗忘与重学。具体地,采用高效的协同训练架构来消除标签噪声的影响,并对每个标签重新校准模型置信度以实现重学。在四个真实数据集上进行的大量实验表明,经由MRR后,COLUR始终优于其他SOTA方法。

原文摘要 · Abstract (English)

Large deep learning models have achieved significant success in various tasks. However, the performance of a model can significantly degrade if it is needed to train on datasets with noisy labels with misleading or ambiguous information. To date, there are limited investigations on how to restore performance when model degradation has been incurred by noisy label data. Inspired by the ``forgetting mechanism'' in neuroscience, which enables accelerating the relearning of correct knowledge by unlearning the wrong knowledge, we propose a robust model restoration and refinement (MRR) framework COLUR, namely Confidence-Oriented Learning, Unlearning and Relearning. Specifically, we implement COLUR with an efficient co-training architecture to unlearn the influence of label noise, and then refine model confidence on each label for relearning. Extensive experiments are conducted on four real datasets and all evaluation results show that COLUR consistently outperforms other SOTA methods after MRR.

模型修复噪声标签协同训练置信度校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。