arXiv:2505.04620cs.CV2025-05被引 26

提出评估多模态大模型通用能力的新框架,打破唯分数论。

On Path to Multimodal Generalist: General-Level and General-Bench

  • 定义5级通用能力标准,量化模型跨模态理解与生成的协同水平。
  • 构建涵盖700+任务、32.5万样本的General-Bench评测集。
  • 揭示当前模型在复杂任务中表现不一致,推动迈向真正通用AI。

多模态大语言模型(MLLM)正快速演进,从早期的专用模型转向多模态通用模型。其能力已从单一模态理解扩展至跨模态生成,并实现细粒度理解与任意模态支持。尽管已有大量评测基准,但高分是否代表更强能力仍存疑。本文提出General-Level评估框架,定义5级通用能力层级,通过“协同性”指标衡量模型在理解与生成、多模态间的一致性。配套构建General-Bench,覆盖超700个任务、325,800个实例,涵盖多种技能、模态与格式。对100多个顶尖MLLM的评测揭示了当前系统的真实能力排名,凸显向人类级人工智能迈进的挑战。本工作为下一代多模态基础模型研究提供坚实评估基础设施,助力通用人工智能(AGI)实现。

原文摘要 · Abstract (English)

The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple modalities, these models have advanced to not only comprehend but also generate across modalities. Their capabilities have expanded from coarse-grained to fine-grained multimodal understanding and from supporting limited modalities to arbitrary ones. While many benchmarks exist to assess MLLMs, a critical question arises: Can we simply assume that higher performance across tasks indicates a stronger MLLM capability, bringing us closer to human-level AI? We argue that the answer is not as straightforward as it seems. This project introduces General-Level, an evaluation framework that defines 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI. At the core of the framework is the concept of Synergy, which measures whether models maintain consistent capabilities across comprehension and generation, and across multiple modalities. To support this evaluation, we present General-Bench, which encompasses a broader spectrum of skills, modalities, formats, and capabilities, including over 700 tasks and 325,800 instances. The evaluation results that involve over 100 existing state-of-the-art MLLMs uncover the capability rankings of generalists, highlighting the challenges in reaching genuine AI. We expect this project to pave the way for future research on next-generation multimodal foundation models, providing a robust infrastructure to accelerate the realization of AGI. Project page: https://generalist.top/

多模态通用模型评测基准AGI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。