arXiv:2409.01560cs.CVcs.AI2024-09中稿 · The 35th British M…被引 1

用积木设计新评测框架,深入检验大模型的分类能力

Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models

  • 以组合积木构建可分解的评测体系,覆盖分类学习到应用全过程
  • 发现大模型在空间关系感知和抽象分类上仍远逊于人类
  • 适合关注模型可解释性与泛化能力的研究者参考

分类是人类基于共性特征组织物体的核心认知能力,在认知科学与计算机视觉中均至关重要。为评估视觉AI模型的分类能力,已有多种从数据集到开放世界场景的代理任务被提出。近年来大型多模态模型(LMMs)在视觉问答、视频时序推理等高层视觉任务中表现优异,得益于先进架构与大规模多模态指令微调。尽管已有整体性基准评估其高层视觉能力,但对最基础的分类能力仍缺乏纯净且深入的量化评估。根据人类认知研究,分类包含类别学习与类别使用两部分。受此启发,我们提出一种基于复合积木的新颖、挑战性强且高效的基准ComBo,提供解耦评估框架,涵盖从学习到应用的完整分类过程。通过对多项任务结果的分析发现,虽然LMMs在学习新类别方面具备可接受的泛化能力,但在细粒度空间关系感知与抽象类别理解方面仍存在明显差距。通过该研究,可为提升LMMs的可解释性与泛化能力提供启示。

原文摘要 · Abstract (English)

Categorization, a core cognitive ability in humans that organizes objects based on common features, is essential to cognitive science as well as computer vision. To evaluate the categorization ability of visual AI models, various proxy tasks on recognition from datasets to open world scenarios have been proposed. Recent development of Large Multimodal Models (LMMs) has demonstrated impressive results in high-level visual tasks, such as visual question answering, video temporal reasoning, etc., utilizing the advanced architectures and large-scale multimodal instruction tuning. Previous researchers have developed holistic benchmarks to measure the high-level visual capability of LMMs, but there is still a lack of pure and in-depth quantitative evaluation of the most fundamental categorization ability. According to the research on human cognitive process, categorization can be seen as including two parts: category learning and category use. Inspired by this, we propose a novel, challenging, and efficient benchmark based on composite blocks, called ComBo, which provides a disentangled evaluation framework and covers the entire categorization process from learning to use. By analyzing the results of multiple evaluation tasks, we find that although LMMs exhibit acceptable generalization ability in learning new categories, there are still gaps compared to humans in many ways, such as fine-grained perception of spatial relationship and abstract category understanding. Through the study of categorization, we can provide inspiration for the further development of LMMs in terms of interpretability and generalization.

多模态模型分类评测认知机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。