用概念融合框架评估并提升视觉语言模型的组合创造力
Probing and Inducing Combinational Creativity in Vision-Language Models
- 提出IEI框架,分解创意生成为识别、提取、推导三步
- 在666张艺术混搭图上验证,顶尖模型理解力超普通人但不及专家
- 将框架融入生成流程,显著提升模型输出的创意质量
人类智能的核心能力之一是将已有概念组合成新想法。近期视觉语言模型(如GPT-4V和DALL·E-3)的进展引发了争议:其输出是否体现真正的组合创造力——即通过整合已有概念生成新颖思想,还是仅是训练数据的复杂模式匹配?受认知科学启发,本文从概念融合视角研究视觉语言模型的组合创造力。我们提出识别-解释-推论(IEI)框架,将创意过程分解为三个层次:识别输入空间、提取共享属性、推导新语义含义。为验证该框架,我们构建了高质量数据集CreativeMashup,包含666张艺术家创作的视觉混搭图,并按IEI框架标注。实验表明,在理解任务中,最优视觉语言模型表现已超越平均人类水平,但仍低于专家级理解;在生成任务中,将IEI框架融入生成流程可显著提升模型输出的创意质量。研究结果既建立了评估人工创造力的理论基础,也为提升视觉语言模型的创造性生成提供了实践指导。
原文摘要 · Abstract (English)
The ability to combine existing concepts into novel ideas stands as a fundamental hallmark of human intelligence. Recent advances in Vision-Language Models (VLMs) like GPT-4V and DALLE-3 have sparked debate about whether their outputs reflect combinational creativity--defined by M. A. Boden (1998) as synthesizing novel ideas through combining existing concepts--or sophisticated pattern matching of training data. Drawing inspiration from cognitive science, we investigate the combinational creativity of VLMs from the lens of concept blending. We propose the Identification-Explanation-Implication (IEI) framework, which decomposes creative processes into three levels: identifying input spaces, extracting shared attributes, and deriving novel semantic implications. To validate this framework, we curate CreativeMashup, a high-quality dataset of 666 artist-generated visual mashups annotated according to the IEI framework. Through extensive experiments, we demonstrate that in comprehension tasks, best VLMs have surpassed average human performance while falling short of expert-level understanding; in generation tasks, incorporating our IEI framework into the generation pipeline significantly enhances the creative quality of VLMs' outputs. Our findings establish both a theoretical foundation for evaluating artificial creativity and practical guidelines for improving creative generation in VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。