arXiv:2503.21851cs.CV2025-03ICCV被引 10

让大模型在开放世界中直接用自然语言识图,突破传统分类限制。

On Large Multimodal Models as Open-World Image Classifiers

  • 用自然语言提示实现无预设类别的图像分类
  • 13个模型在10个数据集上验证,暴露细粒度识别短板
  • 适合研究开放世界视觉理解与提示工程的学者

传统图像分类依赖预定义类别列表。相比之下,大型多模态模型(LMMs)可通过自然语言直接分类图像(如回答‘图像中的主要物体是什么?’),跳过这一限制。然而,现有对LMM分类性能的研究大多局限于封闭世界设定,类别固定。本文填补这一空白,全面评估LMM在真正开放世界中的表现。我们首先形式化该任务,提出评估协议,定义多种指标以衡量预测类别与真实类别的对齐程度。随后在10个基准数据集上评估13个模型,涵盖典型、非典型、细粒度及极细粒度类别,揭示了LMM在该任务中的挑战。基于所提指标的进一步分析揭示了模型错误类型,凸显粒度与细粒度能力不足的问题,并表明定制提示与推理可缓解这些缺陷。

原文摘要 · Abstract (English)

Traditional image classification requires a predefined list of semantic categories. In contrast, Large Multimodal Models (LMMs) can sidestep this requirement by classifying images directly using natural language (e.g., answering the prompt "What is the main object in the image?"). Despite this remarkable capability, most existing studies on LMM classification performance are surprisingly limited in scope, often assuming a closed-world setting with a predefined set of categories. In this work, we address this gap by thoroughly evaluating LMM classification performance in a truly open-world setting. We first formalize the task and introduce an evaluation protocol, defining various metrics to assess the alignment between predicted and ground truth classes. We then evaluate 13 models across 10 benchmarks, encompassing prototypical, non-prototypical, fine-grained, and very fine-grained classes, demonstrating the challenges LMMs face in this task. Further analyses based on the proposed metrics reveal the types of errors LMMs make, highlighting challenges related to granularity and fine-grained capabilities, showing how tailored prompting and reasoning can alleviate them.

多模态模型开放世界图像分类提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。