arXiv:2503.14478cs.CV2025-03ICCV被引 22

评测多模态大模型在真实图像任务中的创造力,发现开源模型表现远逊于闭源模型。

Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM

  • 构建765个图像相关创意任务的多模态评测集,每题设具体评估标准。
  • 开源多模态大模型在创意任务中显著落后于闭源模型,差距明显。
  • 视觉微调可能削弱基础语言模型的创造性,提示训练策略需优化。

创造力是智能的核心能力,涉及在多样情境中生成新颖且恰当解决方案的能力。尽管大型语言模型(LLMs)已在创造性方面得到广泛评估,但对多模态大语言模型(MLLMs)在该领域的评测仍基本空白。为此,我们提出Creation-MMBench,一个专为评估MLLM在真实世界图像任务中创造力而设计的多模态基准。该基准包含765个测试用例,覆盖51个细粒度任务。为确保评估严谨性,我们为每个测试用例定义了具体评价标准,指导对响应质量及与视觉输入的事实一致性评估。实验结果表明,当前开源的MLLM在创意任务中显著弱于闭源模型。此外,分析显示视觉微调可能损害基础语言模型的创造性能力。Creation-MMBench为提升MLLM创造力提供了重要洞察,并为未来多模态生成智能的发展奠定基础。完整数据与评估代码已公开于https://github.com/open-compass/Creation-MMBench。

原文摘要 · Abstract (English)

Creativity is a fundamental aspect of intelligence, involving the ability to generate novel and appropriate solutions across diverse contexts. While Large Language Models (LLMs) have been extensively evaluated for their creative capabilities, the assessment of Multimodal Large Language Models (MLLMs) in this domain remains largely unexplored. To address this gap, we introduce Creation-MMBench, a multimodal benchmark specifically designed to evaluate the creative capabilities of MLLMs in real-world, image-based tasks. The benchmark comprises 765 test cases spanning 51 fine-grained tasks. To ensure rigorous evaluation, we define instance-specific evaluation criteria for each test case, guiding the assessment of both general response quality and factual consistency with visual inputs. Experimental results reveal that current open-source MLLMs significantly underperform compared to proprietary models in creative tasks. Furthermore, our analysis demonstrates that visual fine-tuning can negatively impact the base LLM's creative abilities. Creation-MMBench provides valuable insights for advancing MLLM creativity and establishes a foundation for future improvements in multimodal generative intelligence. Full data and evaluation code is released on https://github.com/open-compass/Creation-MMBench.

多模态创造力评测MLLM基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。