arXiv:2607.23865cs.LG2026-07

用自监督特征的局部平滑性,解决极端噪声标签下的模型训练难题。

XMix: Combating Extremely Noisy Labels via Local Smoothness in Self-Supervised Feature Space

论文配图:XMix: Combating Extremely Noisy Labels via Local Smoothness in Self-Supervised Feature Space
图 1 · 摘自论文原文
  • 基于自监督特征邻域的局部平滑性估计噪声率
  • 利用邻近样本识别更多干净数据并平衡类别选择
  • 在半监督阶段生成更可靠的伪标签,适合高噪声场景

监督深度学习依赖大规模精确标注数据集,但噪声标注普遍存在,高噪声水平下会严重损害性能。现有先进方法通过样本选择策略利用记忆效应筛选干净数据用于半监督学习,但在极端噪声、类别不平衡情况下表现不佳,且需精细调参或先验噪声信息。为此,我们提出XMix框架,利用自监督特征空间中的局部平滑性系统性增强样本选择全过程,不依赖可能被污染的标签。首先,通过自监督特征邻居的最大似然估计噪声率;其次,利用邻居识别额外干净样本,并在样本选择中确保类别均衡;最后,在半监督学习阶段,使用邻近样本生成更可靠的伪标签。实验表明,XMix在极端噪声环境下显著优于现有方法,并在标准噪声标签学习基准上保持优异性能。

原文摘要 · Abstract (English)

Supervised deep learning models rely on large, accurately labeled datasets, yet noisy annotations are often unavoidable and can severely degrade performance under high noise levels. Recent state-of-the-art methods tackle this by using sample selection strategies that exploit the memorization effect to filter out clean data for semi-supervised learning. However, these methods struggle with extreme noise, class imbalance, and require careful tuning or prior noise knowledge. To address these limitations, we propose XMix, a novel framework that leverages local smoothness in the self-supervised feature space to systematically enhance all stages of the sample selection process, without dependence on potentially corrupted labels. First, XMix estimates the noise rate using maximum likelihood among self-supervised feature neighbors. Second, these neighbors then help identify additional clean samples and ensure balanced selection across classes during sample selection. Finally, in the semi-supervised learning phase, XMix uses neighboring samples to generate more reliable pseudo-labels. Our empirical results show that XMix substantially outperforms existing methods in extremely noisy environments and maintains superior performance in standard LNL benchmarks.

噪声标签自监督半监督样本选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。