arXiv:2410.15772cs.LG2024-10综述被引 6

将错误标签检测视为探测模型,提出通用框架与实证基准。

Mislabeled examples detection viewed as probing machine learning models: concepts, survey and extensive benchmark

  • 用四个模块化组件构建统一框架,适配各类分类器
  • 在真实和人工噪声数据上测试,发现现有方法有明显局限
  • 开源工具包支持深度学习与表格数据模型的错误标签探测

真实世界机器学习数据集中普遍存在错误标签,亟需自动检测技术。我们发现多数错误标签检测方法可视为基于若干核心原则对已训练模型的探测。为此,我们形式化一个仅由4个构建模块参数化的模块化框架,并提供一个Python库实现这些原则。重点聚焦于不依赖分类器的方法,强调将为深度模型设计的方法适配到非深度的表格数据分类器。我们在多种任务中对现有方法进行了基准测试,涵盖人工生成的完全随机(NCAR)和更真实的非随机(NNAR)标注噪声。该基准揭示了现有方法在此设定下的新见解与局限性。

原文摘要 · Abstract (English)

Mislabeled examples are ubiquitous in real-world machine learning datasets, advocating the development of techniques for automatic detection. We show that most mislabeled detection methods can be viewed as probing trained machine learning models using a few core principles. We formalize a modular framework that encompasses these methods, parameterized by only 4 building blocks, as well as a Python library that demonstrates that these principles can actually be implemented. The focus is on classifier-agnostic concepts, with an emphasis on adapting methods developed for deep learning models to non-deep classifiers for tabular data. We benchmark existing methods on (artificial) Completely At Random (NCAR) as well as (realistic) Not At Random (NNAR) labeling noise from a variety of tasks with imperfect labeling rules. This benchmark provides new insights as well as limitations of existing methods in this setup.

错误标签检测模型探测数据质量基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。