提出无标注关联性评测基准,揭示多模态大模型在联想能力上的严重短板。
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
- 构建无需人工标注的关联任务基准,基于形容词与动词语义概念设计
- 开源模型在关联任务中表现不佳,GPT-4V仍显著落后于人类水平
- 适合关注多模态模型认知局限与记忆机制的研究者使用
多模态大语言模型(MLLMs)展现出强大能力,但相比人类智能仍存在诸多缺陷,例如幻觉问题。为推动研究进展,社区致力于构建更大型、更复杂的基准测试。本文提出评测一个常被忽视的基础智能——联想能力,即人类将观察与已有经验记忆相联系的能力。为此,我们定义了联想任务,并基于形容词和动词语义概念构建标准基准。不同于耗时的数据标注,我们提出一种便捷的无标注构建方法,将通用数据集转换为联想任务数据;同时设计严格的去噪流程以消除原始数据中的混淆信息。基于该数据集,建立三类联想任务:单步、同步与异步联想。我们系统评估了多种模型在零样本下的联想能力,涵盖三种不同记忆策略、开源与闭源模型、前沿的混合专家(MoE)模型以及人类专家参与情况。结果表明,当前开源模型在联想任务中表现普遍较差,即使最先进的GPT-4V(vision)也与人类存在显著差距。我们认为该基准将为未来多模态大模型研究提供重要方向。数据与代码已公开:https://mvig-rhos.com/llm_inception。
原文摘要 · Abstract (English)
Multi-modal Large Language Models (MLLMs) have exhibited impressive capability. However, recently many deficiencies of MLLMs have been found compared to human intelligence, $\textit{e.g.}$, hallucination. To drive the MLLMs study, the community dedicated efforts to building larger benchmarks with complex tasks. In this paper, we propose benchmarking an essential but usually overlooked intelligence: $\textbf{association}$, a human's basic capability to link observation and prior practice memory. To comprehensively investigate MLLM's performance on the association, we formulate the association task and devise a standard benchmark based on adjective and verb semantic concepts. Instead of costly data annotation and curation, we propose a convenient $\textbf{annotation-free}$ construction method transforming the general dataset for our association tasks. Simultaneously, we devise a rigorous data refinement process to eliminate confusion in the raw dataset. Building on this database, we establish three levels of association tasks: single-step, synchronous, and asynchronous associations. Moreover, we conduct a comprehensive investigation into the MLLMs' zero-shot association capabilities, addressing multiple dimensions, including three distinct memory strategies, both open-source and closed-source MLLMs, cutting-edge Mixture-of-Experts (MoE) models, and the involvement of human experts. Our systematic investigation shows that current open-source MLLMs consistently exhibit poor capability in our association tasks, even the currently state-of-the-art GPT-4V(vision) also has a significant gap compared to humans. We believe our benchmark would pave the way for future MLLM studies. $\textit{Our data and code are available at:}$ https://mvig-rhos.com/llm_inception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。