揭示大模型在多模态评测中依赖语言能力而非视觉推理的真相
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
- 设计新评测协议分离语言与多模态贡献,诊断知识来源
- 发现50%错误源于模型知识不足,无视觉输入也能高分
- 提出知识增强方案,部分数据集性能提升60%
多模态大模型(MLLM)的快速发展伴随各类评测基准的兴起,但这些评测真正衡量的是多模态推理能力,还是仅依赖底层大语言模型(LLM)的先验知识仍不清晰。本文系统研究了LLM主干在MLLM评测中的作用,重点关注两个方面:当前基准是否真正评估多模态推理,以及LLM先验知识对性能的影响。我们提出一种改进的评测协议以解耦LLM主干与多模态融合的贡献,并开发自动知识识别技术,诊断模型对多模态问题所需知识的掌握程度。研究涵盖四个不同多模态基准和八种先进MLLM。关键发现显示,某些基准即使缺少视觉输入也能获得高分,且高达50%的错误可归因于LLM主干世界知识不足,表明性能严重依赖语言能力。为弥补知识缺陷,我们提出知识增强流水线,在特定数据集上实现最高60%的性能提升,整体表现接近4倍增长。本工作揭示了LLM主干在多模态模型中的核心作用,强调需采用更精细的评测方法。
原文摘要 · Abstract (English)
The rapid advancement of Multimodal Large Language Models (MLLMs) has been accompanied by the development of various benchmarks to evaluate their capabilities. However, the true nature of these evaluations and the extent to which they assess multimodal reasoning versus merely leveraging the underlying Large Language Model (LLM) backbone remain unclear. This paper presents a comprehensive investigation into the role of LLM backbones in MLLM evaluation, focusing on two critical aspects: the degree to which current benchmarks truly assess multimodal reasoning and the influence of LLM prior knowledge on performance. Specifically, we introduce a modified evaluation protocol to disentangle the contributions of the LLM backbone from multimodal integration, and an automatic knowledge identification technique for diagnosing whether LLMs equip the necessary knowledge for corresponding multimodal questions. Our study encompasses four diverse MLLM benchmarks and eight state-of-the-art MLLMs. Key findings reveal that some benchmarks allow high performance even without visual inputs and up to 50\% of error rates can be attributed to insufficient world knowledge in the LLM backbone, indicating a heavy reliance on language capabilities. To address knowledge deficiencies, we propose a knowledge augmentation pipeline that achieves significant performance gains, with improvements of up to 60\% on certain datasets, resulting in a approximately 4x increase in performance. Our work provides crucial insights into the role of the LLM backbone in MLLMs, and highlights the need for more nuanced benchmarking approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。