arXiv:2411.14062cs.CVcs.AI2024-11被引 4

自动评估图文模型生成描述能力,发现多数模型表现不佳。

MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective

  • 用图像生成文本再反向生成图像,自动化评估模型描述能力
  • 在13种图像模式上测试50+模型,发现多数优秀模型无法准确描述图像
  • 适合研究图文生成、模型评估的学者使用

大型多模态模型(LMMs)表现出强大能力,但现有评测基准多聚焦特定领域内的图像理解,且构建成本高、回答简短,难以评估模型生成详细图像描述的能力。为此,我们提出MMGenBench-Pipeline,一种简单且全自动的评估流程:从输入图像生成文本描述,利用文本-图像生成模型生成辅助图像,并比较原始与生成图像的一致性。为确保有效性,设计了MMGenBench-Test(覆盖13种图像模式)和MMGenBench-Domain(聚焦生成性能)。对超过50个主流LMMs的全面评估表明该流程和基准具有有效性和可靠性。观察发现,许多在现有基准中表现优异的模型,在图像理解与描述任务上仍存在明显不足,揭示当前模型仍有巨大优化空间。同时,该管道仅需图像输入即可高效评估模型在多种领域的表现。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) demonstrate impressive capabilities. However, current benchmarks predominantly focus on image comprehension in specific domains, and these benchmarks are labor-intensive to construct. Moreover, their answers tend to be brief, making it difficult to assess the ability of LMMs to generate detailed descriptions of images. To address these limitations, we propose the MMGenBench-Pipeline, a straightforward and fully automated evaluation pipeline. This involves generating textual descriptions from input images, using these descriptions to create auxiliary images via text-to-image generative models, and then comparing the original and generated images. Furthermore, to ensure the effectiveness of MMGenBench-Pipeline, we design MMGenBench-Test, evaluating LMMs across 13 distinct image patterns, and MMGenBench-Domain, focusing on generative image performance. A thorough evaluation involving over 50 popular LMMs demonstrates the effectiveness and reliability of both the pipeline and benchmark. Our observations indicate that numerous LMMs excelling in existing benchmarks fail to adequately complete the basic tasks related to image understanding and description. This finding highlights the substantial potential for performance improvement in current LMMs and suggests avenues for future model optimization. Concurrently, MMGenBench-Pipeline can efficiently assess the performance of LMMs across diverse domains using only image inputs.

多模态评估图像生成自动评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。