arXiv:2411.14725cs.CVcs.CL2024-11被引 2

构建统一评估框架,精准测评多模态模型视觉感知能力

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens

  • 设计六维感知能力统一评测基准,覆盖计数、OCR等场景
  • 发现顶尖开源与闭源模型间存在显著性能差距,且存在能力冲突现象
  • 揭示数据混合比例和语言模型规模是能力冲突主因,适合模型优化研究者参考

随着多模态大语言模型(MLLMs)快速发展,严格的评估成为关键。本文聚焦于统一且稳健的视觉感知能力评估——这是MLLM的基础技能。现有感知基准各自侧重不同题型、领域和评价指标,导致评估结果差异显著,难以全面评估感知能力。为此,我们提出AbilityLens,一个涵盖六项核心感知能力(从计数、OCR到结构化数据理解)的统一基准,关注准确率与稳定性,每项能力均包含多样题型、领域和度量方式。借助AbilityLens,我们:(1)识别主流MLLM的优劣势,揭示稳定性模式并发现先进开源与闭源模型间的显著性能差距;(2)发现训练过程中的能力冲突与早期收敛现象;(3)证实能力冲突主要源于数据混合比例与语言模型规模;(4)讨论微调与模型融合等策略对缓解冲突的有效性。基准与在线排行榜已公开于https://github.com/Chenfeng1271/AbilityLens。

原文摘要 · Abstract (English)

As multimodal large language models (MLLMs) advance rapidly, rigorous evaluation has become essential, providing further guidance for their development. In this work, we focus on a unified and robust evaluation of \textbf{vision perception} abilities, the foundational skill of MLLMs. We find that existing perception benchmarks, each focusing on different question types, domains, and evaluation metrics, introduce significant evaluation variance, complicating comprehensive assessments of perception abilities when relying on any single benchmark. To address this, we introduce \textbf{AbilityLens}, a unified benchmark designed to evaluate MLLMs in six key perception abilities (ranging from counting, OCR, to understanding structural data), focusing on both accuracy and stability, with each ability encompassing diverse types of questions, domains, and metrics. With the assistance of AbilityLens, we: (1) identify the strengths and weaknesses of current main-stream MLLMs, highlighting stability patterns and revealing a notable performance gap between state-of-the-art open-source and closed-source models; (2) uncover interesting ability conflict and early convergence phenomena during MLLM training; (3) reveal the primary reason of ability conflict is data mixing ratio and LLM model size; and (4) discuss the effectiveness of some straightforward strategies \eg, fine-tuning and model merging, to solve the ability conflict. The benchmark and online leaderboard is released in https://github.com/Chenfeng1271/AbilityLens.

多模态模型评估基准感知能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。