构建首个统一多模态理解生成评估基准,填补评测空白。
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
- 设计标准化任务集覆盖10类30子任务,实现公平对比。
- 新增5项混合模态生成任务,检验跨模态推理能力。
- 评测12款主流模型,揭示当前统一模型性能短板。
现有多模态大模型评测基准在评估统一型多模态大模型(U-MLLMs)时面临两大挑战:一是缺乏传统任务的标准化评测,导致研究间比较不一致;二是缺少混合模态生成任务评测,难以评估跨模态推理能力。为此,我们提出一个全面的评估框架,包含三方面:1)标准化传统任务评测,从12个数据集采样,覆盖10类任务共30个子任务,确保研究间可比性;2)统一任务评估,引入5项新任务,测试图像编辑、带图像生成的常识问答与几何推理等多模态推理能力;3)全面模型评测,涵盖Janus-Pro、EMU3、VILA-U、Gemini2-flash等12款领先统一模型,以及Claude-3.5-Sonnet等专用理解模型和DALL-E-3等生成模型。结果表明,现有U-MLLMs在混合模态任务上存在显著性能差距,亟需更强大的模型以应对复杂跨模态任务。代码与评测数据见https://mme-unify.github.io/。
原文摘要 · Abstract (English)
Existing MLLM benchmarks face significant challenges in evaluating Unified MLLMs (U-MLLMs) due to: 1) lack of standardized benchmarks for traditional tasks, leading to inconsistent comparisons; 2) absence of benchmarks for mixed-modality generation, which fails to assess multimodal reasoning capabilities. We present a comprehensive evaluation framework designed to systematically assess U-MLLMs. Our benchmark includes: Standardized Traditional Task Evaluation. We sample from 12 datasets, covering 10 tasks with 30 subtasks, ensuring consistent and fair comparisons across studies." 2. Unified Task Assessment. We introduce five novel tasks testing multimodal reasoning, including image editing, commonsense QA with image generation, and geometric reasoning. 3. Comprehensive Model Benchmarking. We evaluate 12 leading U-MLLMs, such as Janus-Pro, EMU3, VILA-U, and Gemini2-flash, alongside specialized understanding (e.g., Claude-3.5-Sonnet) and generation models (e.g., DALL-E-3). Our findings reveal substantial performance gaps in existing U-MLLMs, highlighting the need for more robust models capable of handling mixed-modality tasks effectively. The code and evaluation data can be found in https://mme-unify.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。