arXiv:2603.09970cs.CL2026-03被引 7

测试大模型联想创造力的新基准,评估其生成独特连接路径的能力。

CREATE: Testing LLMs for Associative Creativity

论文配图:CREATE: Testing LLMs for Associative Creativity
图 1 · 摘自论文原文
  • 通过生成概念间的高特异性、多样化连接路径来评估创造力。
  • 顶尖模型表现更优,但因搜索空间巨大,基准难达饱和。
  • 思考型模型未必更有效,现有创意提示方法提升有限。

创造力的核心是联想推理:在概念间建立新颖而有意义的联系。我们提出 CREATE 基准,用于评估模型在参数知识中生成概念连接路径的创造性联想能力。路径需具备高特异性(连接的独特性与紧密度)和高多样性(与其他路径差异性),模型因生成更多强且多样的路径而得分更高。该任务模拟真实创造力任务(如假设生成)的复杂性,拥有极大搜索空间,但支持客观评分的大规模数据收集。对前沿模型的评估显示,最强模型在创造效用上优于其他模型,且答案多样性与搜索复杂性导致基准难以饱和。此外,结果表明,具有高令牌预算的思考型模型在本任务中并不总是更高效;近期创意提示方法仅带来有限改进。CREATE 为提升模型联想创造力的新方法提供了实验平台。

原文摘要 · Abstract (English)

A key component of creativity is associative reasoning: the ability to draw novel yet meaningful connections between concepts. We introduce CREATE, a benchmark designed to evaluate models' capacity for creative associative reasoning. CREATE requires models to generate sets of paths connecting concepts in a model's parametric knowledge. Paths should have high specificity (distinctiveness and closeness of the concept connection) and high diversity (dissimilarity from other paths), and models are scored more highly if they produce a larger set of strong, diverse paths. This task shares demands of real creativity tasks like hypothesis generation, including an extremely large search space, but enables collection of a sizable benchmark with objective answer grading. Evaluation of frontier models shows that the strongest models achieve higher creative utility than others, with the high multiplicity of answers and complexity of the search making benchmark saturation difficult to achieve. Furthermore, our results illustrate that thinking models are not always more effective on our task, even with high token budgets. Recent approaches for creative prompting give some but limited additional improvement. CREATE provides a sandbox for developing new methods to improve models' capacity for associative creativity.

大模型创造力基准测试联想推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。