用合成图像修正真实数据中的语义误标,提升分类准确率。
Noisy Label Refinement with Semantically Reliable Synthetic Images
- 用高质量合成图像作可靠标签参考,定位真实数据中的误标样本。
- 在70%语义噪声下,CIFAR-10准确率提升30%,ImageNet-100提升24%。
- 可与现有抗噪方法叠加使用,适合数据质量差的场景。
图像分类数据集中存在的语义噪声(视觉相似类别常被误标)对传统监督学习构成重大挑战。本文探索利用先进文本到图像模型生成的高质量合成图像来解决该问题。尽管这些合成图像具有可靠标签,但其直接用于训练受限于领域差异和多样性约束。不同于传统方法,我们提出一种新策略:将合成图像作为可靠参照点,识别并修正噪声数据集中的误标样本。在多个基准数据集上的大量实验表明,该方法在各种噪声条件下显著提升分类准确率,尤其在高难度语义噪声场景下表现突出。此外,由于该方法与现有抗噪学习技术正交,与当前最先进的抗噪训练方法结合后性能更优:在70%语义噪声下,CIFAR-10准确率提升30%,CIFAR-100提升11%;在真实世界噪声条件下,ImageNet-100准确率提升24%。
原文摘要 · Abstract (English)
Semantic noise in image classification datasets, where visually similar categories are frequently mislabeled, poses a significant challenge to conventional supervised learning approaches. In this paper, we explore the potential of using synthetic images generated by advanced text-to-image models to address this issue. Although these high-quality synthetic images come with reliable labels, their direct application in training is limited by domain gaps and diversity constraints. Unlike conventional approaches, we propose a novel method that leverages synthetic images as reliable reference points to identify and correct mislabeled samples in noisy datasets. Extensive experiments across multiple benchmark datasets show that our approach significantly improves classification accuracy under various noise conditions, especially in challenging scenarios with semantic label noise. Additionally, since our method is orthogonal to existing noise-robust learning techniques, when combined with state-of-the-art noise-robust training methods, it achieves superior performance, improving accuracy by 30% on CIFAR-10 and by 11% on CIFAR-100 under 70% semantic noise, and by 24% on ImageNet-100 under real-world noise conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。