arXiv:2504.13120cs.CVcs.AI2025-04被引 7

用概念融合框架评估并提升视觉语言模型的组合创造力

Probing and Inducing Combinational Creativity in Vision-Language Models

  • 提出IEI框架,分解创意生成为识别、提取、推导三步
  • 在666张艺术混搭图上验证,顶尖模型理解力超普通人但不及专家
  • 将框架融入生成流程,显著提升模型输出的创意质量

人类智能的核心能力之一是将已有概念组合成新想法。近期视觉语言模型(如GPT-4V和DALL·E-3)的进展引发了争议:其输出是否体现真正的组合创造力——即通过整合已有概念生成新颖思想,还是仅是训练数据的复杂模式匹配?受认知科学启发,本文从概念融合视角研究视觉语言模型的组合创造力。我们提出识别-解释-推论(IEI)框架,将创意过程分解为三个层次:识别输入空间、提取共享属性、推导新语义含义。为验证该框架,我们构建了高质量数据集CreativeMashup,包含666张艺术家创作的视觉混搭图,并按IEI框架标注。实验表明,在理解任务中,最优视觉语言模型表现已超越平均人类水平,但仍低于专家级理解;在生成任务中,将IEI框架融入生成流程可显著提升模型输出的创意质量。研究结果既建立了评估人工创造力的理论基础,也为提升视觉语言模型的创造性生成提供了实践指导。

原文摘要 · Abstract (English)

The ability to combine existing concepts into novel ideas stands as a fundamental hallmark of human intelligence. Recent advances in Vision-Language Models (VLMs) like GPT-4V and DALLE-3 have sparked debate about whether their outputs reflect combinational creativity--defined by M. A. Boden (1998) as synthesizing novel ideas through combining existing concepts--or sophisticated pattern matching of training data. Drawing inspiration from cognitive science, we investigate the combinational creativity of VLMs from the lens of concept blending. We propose the Identification-Explanation-Implication (IEI) framework, which decomposes creative processes into three levels: identifying input spaces, extracting shared attributes, and deriving novel semantic implications. To validate this framework, we curate CreativeMashup, a high-quality dataset of 666 artist-generated visual mashups annotated according to the IEI framework. Through extensive experiments, we demonstrate that in comprehension tasks, best VLMs have surpassed average human performance while falling short of expert-level understanding; in generation tasks, incorporating our IEI framework into the generation pipeline significantly enhances the creative quality of VLMs' outputs. Our findings establish both a theoretical foundation for evaluating artificial creativity and practical guidelines for improving creative generation in VLMs.

视觉语言模型组合创造力概念融合生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。