arXiv:2505.16149cs.CVcs.AI2025-05被引 4

用视觉语言模型发现图像分类数据集中的漏标问题

When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification

  • 融合多模型与人工标注,识别数据集中噪声标签和缺失标签
  • 在6个基准测试集上显著提升标签准确性,接近人类判断水平
  • 适合关注数据质量、模型公平评估的研究者使用

图像分类基准数据集如CIFAR、MNIST和ImageNet是模型评估的关键工具。然而,尽管经过清理,这些数据集仍普遍存在标签噪声,并因图像中同时包含多个类别而存在缺失标签问题,导致模型比较失真、评价不公平。现有标签清洗方法主要针对噪声标签,对缺失标签问题关注不足。为此,我们提出综合性框架REVEAL,结合先进预训练视觉语言模型(如LLaVA、BLIP、Janus、Qwen)与机器/人工标签清理方法(如Docta、Cleanlab、MTurk),系统解决常见图像分类测试集中的噪声与漏标问题。REVEAL检测潜在噪声与遗漏,通过置信度加权与共识过滤聚合预测结果,生成带概率的软标签。经人工验证,该方法显著提升6个基准测试集质量,高度契合人类判断,实现更准确、有意义的图像分类比较。

原文摘要 · Abstract (English)

Image classification benchmark datasets such as CIFAR, MNIST, and ImageNet serve as critical tools for model evaluation. However, despite the cleaning efforts, these datasets still suffer from pervasive noisy labels and often contain missing labels due to the co-existing image pattern where multiple classes appear in an image sample. This results in misleading model comparisons and unfair evaluations. Existing label cleaning methods focus primarily on noisy labels, but the issue of missing labels remains largely overlooked. Motivated by these challenges, we present a comprehensive framework named REVEAL, integrating state-of-the-art pre-trained vision-language models (e.g., LLaVA, BLIP, Janus, Qwen) with advanced machine/human label curation methods (e.g., Docta, Cleanlab, MTurk), to systematically address both noisy labels and missing label detection in widely-used image classification test sets. REVEAL detects potential noisy labels and omissions, aggregates predictions from various methods, and refines label accuracy through confidence-informed predictions and consensus-based filtering. Additionally, we provide a thorough analysis of state-of-the-art vision-language models and pre-trained image classifiers, highlighting their strengths and limitations within the context of dataset renovation by revealing 10 observations. Our method effectively reveals missing labels from public datasets and provides soft-labeled results with likelihoods. Through human verifications, REVEAL significantly improves the quality of 6 benchmark test sets, highly aligning to human judgments and enabling more accurate and meaningful comparisons in image classification.

数据清洗视觉语言模型标签质量图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。