arXiv:2606.11699cs.LG2026-06

提出端到端数据框架,自动发现并修正错误标签

A Data-Centric Framework for Detecting and Correcting Corrupted Labels

论文配图:A Data-Centric Framework for Detecting and Correcting Corrupted Labels
图 1 · 摘自论文原文
  • 结合局部与全局数据关系识别噪声样本
  • 通过特征和噪声标签估计最可能的正确标签
  • 在多数据集上显著提升标签修正精度与任务性能

机器学习与深度学习模型的性能高度依赖训练数据质量。然而,真实世界数据集常因标签噪声而受损,严重降低模型准确率与可靠性。为此,我们提出Relabeler——一个端到端的数据中心框架,用于检测与修正错误标签。在检测阶段,Relabeler联合利用数据实例间的局部与全局关系,识别潜在噪声样本;在修正阶段,基于输入特征与观测到的噪声标签,估计每个样本最可能的干净标签。在多个数据集、噪声类型与噪声率下的大量实验表明,Relabeler持续优于现有先进基线方法,在标签修正精度上最高提升58%,下游任务性能提升6%。

原文摘要 · Abstract (English)

The performance of machine learning and deep learning models largely depends on the quality of the training data. However, the quality of the real-world datasets is often compromised by noisy labels, which can substantially degrade model accuracy and reliability. To address this challenge, we propose Relabeler, an end-to-end data-centric framework for detecting and correcting corrupted labels. For corrupted label detection, Relabeler jointly leverages both local and global relationships among data instances to identify potentially noisy samples. After detecting suspicious instances, Relabeler further performs label correction by estimating the most probable clean label for each instance based on both its input features and observed noisy label. Extensive experiments across multiple datasets, noise types, and noise rates demonstrate that Relabeler consistently outperforms state-of-the-art baselines, achieving up to 58% improvement in label correction precision and 6% improvement in downstream task performance.

数据清洗标签纠错数据质量机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。