评测大模型跨概念理解能力,发现其创意解码仍有巨大差距。
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding

- 构建C4框架,用成语联想路径测试模型跨概念推理能力。
- 十款模型最高准确率仅50.7%,开源模型表现明显偏低。
- 显式约束能显著提升性能,但提示解释效果有限。
大模型在设计、教育与人机协作中的创造能力至关重要,但因缺乏明确目标与奖励信号,评估困难。跨概念理解是接受性创造力的核心,使感知者从非显性但有意义的概念关联中还原意图。本文将题目构建定义为跨概念编码,模型推断定义为跨概念解码。提出受认知启发的C4评估框架,基于成语(Chengyu)进行跨概念创造性评估。其编码部分通过人工标注并经第三方审核的跨概念网络,将目标槽位映射到可图像化的替代概念,沿桥接路径生成,支持批量构造,难度由桥接数量和深度控制,答案精确可验证。基于此,构建了包含184个合成项和37个来自网络的人类创作成语图式的C4-Eval数据集。对收集的成语图式,人工构建并审核其概念关系、桥接路径与推理过程。每个条目在五种任务设置下实例化,共产生884个主要答案恢复案例。在十款评估的大模型中,最强闭源模型达到50.7%和48.0%的主准确率,而开源模型仍显著更低。候选约束条件显著提升准确率,但桥接提示和解释请求仅带来适度改善。结果揭示当前大模型在通过跨概念关系解码创意编码信息方面存在明显短板。代码见附录。
原文摘要 · Abstract (English)
Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core cognitive capacity underlying receptive creativity. It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations. We operationalize item construction as cross-concept encoding and model inference as cross-concept decoding. We introduce C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity. Its encoding component maps target slots to imageable substitute concepts along bridge paths in a manually annotated and third-party-reviewed cross-concept network, enabling batch generation with explicit structure, difficulty indexed by bridge count and depth, and exact answers. Using this framework, we instantiate the C4 Evaluation Set (C4-Eval), comprising 184 synthetic items and 37 human-created cross-concept chengyu figures collected from online sources. We manually construct and review cross-concept relations, bridge paths, and reasoning processes for the collected figures. Each C4-Eval item is instantiated in five task settings, yielding 884 primary answer-recovery cases. Across ten evaluated MLLMs, the strongest closed models reach 50.7% and 48.0% primary accuracy, while open-source models remain substantially lower. Candidate constraints improve accuracy sharply, but bridge hints and explanation requests provide only modest gains. These results expose a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations. The code is in the supplementary material.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。