arXiv:2603.06578cs.CV2026-03被引 2

MLLM图像分类性能受标注质量和评估协议影响,修正后显著提升。

Multimodal Large Language Models as Image Classifiers

  • 改进评估协议,修复输出过滤、干扰项和映射错误问题
  • 在ReGT数据集上,修正标签使准确率最高提升10.8%
  • 适合关注标注质量与模型真实能力的研究者

多模态大语言模型(MLLM)的图像分类性能高度依赖于评估协议和真实标签质量。现有研究对比MLLM与监督模型时结论矛盾,根源在于评估方法或高估或低估性能。我们识别并修正了常见评估中的关键问题:模型输出超出类别列表被丢弃、弱多项选择干扰项导致结果虚高、开放世界设置因输出映射不佳而表现差。此外,量化发现批量大小、图像排序和文本编码器选择对准确率有显著影响。在ReGT(对ImageNet-1k 625个类别的多标签重标注)上评估显示,修正标签后MLLM性能最高提升10.8%,大幅缩小与监督模型的差距。多数报告中MLLM分类性能不足实为噪声标签与评估缺陷所致,而非模型本身缺陷。对监督信号依赖少的模型更敏感于标注质量。最后,在可控案例研究中,人工标注者在约50%的疑难样本中采纳或确认了MLLM预测,表明其可用于大规模数据集构建。本工作属于Aiming for Perfect ImageNet-1k项目,详见https://klarajanouskova.github.io/ImageNet/。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLM) classification performance depends critically on evaluation protocol and ground truth quality. Studies comparing MLLMs with supervised and vision-language models report conflicting conclusions, and we show these conflicts stem from protocols that either inflate or underestimate performance. Across the most common evaluation protocols, we identify and fix key issues: model outputs that fall outside the provided class list and are discarded, inflated results from weak multiple-choice distractors, and an open-world setting that underperforms only due to poor output mapping. We additionally quantify the impact of commonly overlooked design choices - batch size, image ordering, and text encoder selection - showing they substantially affect accuracy. Evaluating on ReGT, our multilabel reannotation of 625 ImageNet-1k classes, reveals that MLLMs benefit most from corrected labels (up to +10.8%), substantially narrowing the perceived gap with supervised models. Much of the reported MLLMs underperformance on classification is thus an artifact of noisy ground truth and flawed evaluation protocol rather than genuine model deficiency. Models less reliant on supervised training signals prove most sensitive to annotation quality. Finally, we show that MLLMs can assist human annotators: in a controlled case study, annotators confirmed or integrated MLLMs predictions in approximately 50% of difficult cases, demonstrating their potential for large-scale dataset curation. This work is part of the Aiming for Perfect ImageNet-1k project, see https://klarajanouskova.github.io/ImageNet/.

多模态图像分类标注质量评估协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。