统一评估多模态模型的理解、生成与编辑能力
UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
- 用自洽评估范式在单次测试中覆盖理解、生成与编辑
- 涵盖13个领域200多个概念,支持细粒度能力拆分
- 适合作为模型开发与评测的基准工具
统一多模态理解与生成在前沿闭源系统中展现出强大能力,但现有评估仍分离进行,分别使用不同数据集测试理解与生成能力。为此,我们提出UmniBench,一个面向统一多模态模型(UMMs)的全方位评测基准。首先,UmniBench可在单一评估流程中同时检验模型的理解、生成与编辑能力;基于人工验证的提示与问答对,利用模型自身来评估其生成与编辑表现,结合理解能力实现闭环评估。该方法简单有效,实现全面测评。其次,基准覆盖13个主要领域及200多个概念,确保对模型能力的充分检验。此外,也支持解耦评估,可单独分析理解、生成与编辑性能。基于此,我们对24个主流模型(含多种UMMs与单能力大模型)进行了评测,旨在为社区提供更全面客观的统一模型评估视角,并推动模型性能持续优化。
原文摘要 · Abstract (English)
Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. However, evaluations of unified multimodal models (UMMs) remain decoupled, assessing their understanding and generation abilities separately with corresponding datasets. To address this, we propose UmniBench, a benchmark tailored for UMMs with omni-dimensional evaluation. First, UmniBench can assess the understanding, generation, and editing ability within a single evaluation process. Based on human-examined prompts and QA pairs, UmniBench leverages UMM itself to evaluate its generation and editing ability with its understanding ability. This simple but effective paradigm allows comprehensive evaluation of UMMs. Second, UmniBench covers 13 major domains and more than 200 concepts, ensuring a thorough inspection of UMMs. Moreover, UmniBench can also decouple and separately evaluate understanding, generation, and editing abilities, providing a fine-grained assessment. Based on UmniBench, we benchmark 24 popular models, including both UMMs and single-ability large models. We hope this benchmark provides a more comprehensive and objective view of unified models and logistical support for improving the performance of the community model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。