评测50+任务的多模态生成能力,发现主流模型在复杂指令上表现差异大。
OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
- 设计57个真实场景任务,分感知与认知两类评估模型能力
- 用现成视觉工具和大模型评判器双模式测试,结果更可信
- 涵盖GPT-4o等主流模型,适合研究多模态生成的开发者参考
近年来,如GPT-4o-Native等大型多模态模型(LMMs)在图像生成通用指令方面表现出色。然而,现有评测基准往往缺乏足够的广度与深度,难以全面评估这些模型的多样化能力。为此,我们提出OmniGenBench,一个全新且全面的基准,旨在系统评估前沿LMMs在感知型与认知型任务上的指令遵循能力。该基准包含57个基于真实场景的子任务,按所需模型能力进行系统分类。为确保评估严谨性,采用双模式协议:对感知类任务使用现成视觉解析工具,对认知类任务则使用强大的大语言模型作为评判器,以衡量生成图像与用户指令的一致性。利用OmniGenBench,我们评估了包括GPT-4o、Gemini-2.0-Flash、Seedream在内的主流生成模型,并提供了深入的性能对比与分析。代码与数据已开源于https://github.com/emilia113/OmniGenBench。
原文摘要 · Abstract (English)
Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation. However, current benchmarks often lack the necessary breadth and depth to fully evaluate the diverse capabilities of these models. To overcome this limitation, we introduce OmniGenBench, a novel and comprehensive benchmark meticulously designed to assess the instruction-following abilities of state-of-the-art LMMs across both perception-centric and cognition-centric dimensions. Our OmniGenBench includes 57 diverse sub-tasks grounded in real-world scenarios, systematically categorized according to the specific model capabilities they demand. For rigorous evaluation, we further employ a dual-mode protocol. This protocol utilizes off-the-shelf visual parsing tools for perception-centric tasks and a powerful LLM-based judger for cognition-centric tasks to assess the alignment between generated images and user instructions. Using OmniGenBench, we evaluate mainstream generative models, including prevalent models like GPT-4o, Gemini-2.0-Flash, and Seedream, and provide in-depth comparisons and analyses of their performance.Code and data are available at https://github.com/emilia113/OmniGenBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。