arXiv:2502.17771cs.LGcs.AI2025-02NeurIPS被引 6

通过对比片段化提升噪声标签回归中的干净样本选择能力

Sample Selection via Contrastive Fragmentation for Noisy Label Regression

  • 将数据转化为互斥且对比的片段对,增强特征表示区分度
  • 在六大数据集上优于14个先进基线,对对称和高斯噪声均鲁棒
  • 提出新指标ERR,更好衡量不同噪声程度下的性能

真实世界回归任务普遍面临噪声标签问题。我们发现,许多实际数据具有标签与特征间连续有序的相关性:相似标签的数据点其特征也相近。为此,我们提出一种名为ConFrag的新方法,将回归数据转化为不相交但具有对比性的片段对,以训练更具区分性的表示,从而提升干净样本的选择能力。ConFrag框架利用邻近片段的混合,通过多个专家特征提取器之间的邻域一致性来识别噪声标签。我们在六个新构建的跨领域基准数据集(包括年龄预测、价格预测、音乐制作年份估计)上进行了广泛实验,并引入误差残差比(Error Residual Ratio, ERR)作为新评估指标,以更好反映不同噪声程度的影响。结果表明,该方法在所有设置下均显著优于14个当前最优基线,在对称噪声和随机高斯噪声下表现稳健。

原文摘要 · Abstract (English)

As with many other problems, real-world regression is plagued by the presence of noisy labels, an inevitable issue that demands our attention. Fortunately, much real-world data often exhibits an intrinsic property of continuously ordered correlations between labels and features, where data points with similar labels are also represented with closely related features. In response, we propose a novel approach named ConFrag, where we collectively model the regression data by transforming them into disjoint yet contrasting fragmentation pairs. This enables the training of more distinctive representations, enhancing the ability to select clean samples. Our ConFrag framework leverages a mixture of neighboring fragments to discern noisy labels through neighborhood agreement among expert feature extractors. We extensively perform experiments on six newly curated benchmark datasets of diverse domains, including age prediction, price prediction, and music production year estimation. We also introduce a metric called Error Residual Ratio (ERR) to better account for varying degrees of label noise. Our approach consistently outperforms fourteen state-of-the-art baselines, being robust against symmetric and random Gaussian label noise.

噪声标签回归样本选择对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。