校准模型能更准确地发现错误标注数据,提升系统可靠性。
Calibration improves detection of mislabeled examples
- 用校准技术优化基础模型,提高误标检测的可信度。
- 校准后检测准确率和鲁棒性显著提升。
- 适合工业场景中数据质量治理与模型优化。
错误标注数据在真实应用中普遍存在,严重影响机器学习系统的性能。有效缓解该问题的方法是识别误标样本并进行过滤或重标注等特殊处理。现有自动误标检测方法通常依赖训练一个基础模型,再通过该模型对每个样本生成信任分数以判断标签是否真实。因此,基础模型的特性至关重要。本文研究了对模型进行校准的影响。实验结果表明,采用校准方法可显著提升误标样本检测的准确率与鲁棒性,为工业应用提供了一种实用而有效的解决方案。
原文摘要 · Abstract (English)
Mislabeled data is a pervasive issue that undermines the performance of machine learning systems in real-world applications. An effective approach to mitigate this problem is to detect mislabeled instances and subject them to special treatment, such as filtering or relabeling. Automatic mislabeling detection methods typically rely on training a base machine learning model and then probing it for each instance to obtain a trust score that each provided label is genuine or incorrect. The properties of this base model are thus of paramount importance. In this paper, we investigate the impact of calibrating this model. Our empirical results show that using calibration methods improves the accuracy and robustness of mislabeled instance detection, providing a practical and effective solution for industrial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。