arXiv:2410.10139cs.CVcs.CL2024-10ICLR被引 39

构建大规模多模态交错理解评测集,提升视觉语言模型评估可靠性

MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models

  • 设计20000条跨领域多模态交错问答,支持图文混合输入输出
  • 提出基于人工标注微调的自动化评分模型,减少评估偏差
  • 覆盖数学、编程、医学等102个子领域,适合评估高级多模态能力

交错式多模态理解与生成(即模型可任意顺序处理和生成图像与文本)已成为多模态学习的关键方向。尽管进展显著,现有评测仍存在数据规模小、覆盖范围窄、评估深度不足等问题,且主流指标成本高或存在偏倚,难以可靠用于实际应用。为此,我们提出MMIE——一个大规模知识密集型多模态交错理解与生成评测基准,专为大视觉语言模型(LVLMs)设计。该基准包含20,000条精心筛选的多模态问题,涵盖3类任务、12个领域及102个子领域,如数学、编程、物理、文学、健康与艺术。支持交错式输入与输出,提供多项选择与开放问答相结合的测评形式,全面评估多种能力。同时,我们提出一种可靠的自动化评估指标,基于人工标注数据微调评分模型,并结合系统性评价标准,以降低偏倚、提升准确性。大量实验表明,该基准与指标能有效全面评估交错式LVLM。我们评估了8个主流LVLM,结果显示即使最佳模型仍有显著提升空间,多数仅达到中等水平。我们公开发布评测集与代码:https://mmie-bench.github.io/。

原文摘要 · Abstract (English)

Interleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancements, the evaluation of this capability remains insufficient. Existing benchmarks suffer from limitations in data scale, scope, and evaluation depth, while current evaluation metrics are often costly or biased, lacking in reliability for practical applications. To address these challenges, we introduce MMIE, a large-scale knowledge-intensive benchmark for evaluating interleaved multimodal comprehension and generation in Large Vision-Language Models (LVLMs). MMIE comprises 20K meticulously curated multimodal queries, spanning 3 categories, 12 fields, and 102 subfields, including mathematics, coding, physics, literature, health, and arts. It supports both interleaved inputs and outputs, offering a mix of multiple-choice and open-ended question formats to evaluate diverse competencies. Moreover, we propose a reliable automated evaluation metric, leveraging a scoring model fine-tuned with human-annotated data and systematic evaluation criteria, aimed at reducing bias and improving evaluation accuracy. Extensive experiments demonstrate the effectiveness of our benchmark and metrics in providing a comprehensive evaluation of interleaved LVLMs. Specifically, we evaluate eight LVLMs, revealing that even the best models show significant room for improvement, with most achieving only moderate results. We believe MMIE will drive further advancements in the development of interleaved LVLMs. We publicly release our benchmark and code in https://mmie-bench.github.io/.

多模态评测基准视觉语言模型自动评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。