arXiv:2509.21043cs.AIcs.LG2025-09被引 8

提出评估大模型创造力的新框架,揭示其能力随规模变化的规律。

Combinatorial Creativity: A New Frontier in Generalization Abilities

  • 用新颖性与实用性双重指标评估模型创造力
  • 发现固定算力下存在最优模型深度与宽度
  • 揭示创意生成中新颖性与可行性间的根本权衡

人工智能系统,特别是大语言模型(LLMs),正越来越多地被用于科学创意生成等创造性任务,这属于训练数据之外的泛化能力,现有理论框架尚无法涵盖。尽管与组合泛化(CG)相似,但组合式创造力(CC)具有开放性特征。由于其开放性,传统以准确性或正确性为标准的评估方式不适用。为此,我们提出一种理论框架与算法任务,通过新颖性与实用性来评估输出。实证发现:(1)首次揭示了大模型创造力的缩放行为;(2)在固定计算预算下,存在最优模型深度与宽度;(3)模型在生成新颖科学构想方面表现优异,但在确保可行性上表现不佳,这一“构想-执行差距”可归因于普遍存在的新颖性-实用性权衡。尽管研究结果在1亿参数量级仍成立,当前前沿模型已达数十亿参数,因此本框架与发现可作为理解并提升超大规模模型创造力的起点,助力缩小人机智能差距。

原文摘要 · Abstract (English)

Artificial intelligence (AI) systems, and Large Language Models (LLMs) in particular, are increasingly employed for creative tasks like scientific idea generation, constituting a form of generalization from training data unaddressed by existing conceptual frameworks. Despite its similarities to compositional generalization (CG), combinatorial creativity (CC) is an open-ended ability. Instead of evaluating for accuracy or correctness against fixed targets, which would contradict the open-ended nature of CC, we propose a theoretical framework and algorithmic task for evaluating outputs by their degrees of novelty and utility. From here, we make several important empirical contributions: (1) We obtain the first insights into the scaling behavior of creativity for LLMs. (2) We discover that, for fixed compute budgets, there exist optimal model depths and widths for creative ability. (3) We find that the ideation-execution gap, whereby LLMs excel at generating novel scientific ideas but struggle to ensure their practical feasibility, may be explained by a more fundamental novelty-utility tradeoff characteristic of creativity algorithms in general. Though our findings persist up to the 100M scale, frontier models today are well into the billions of parameters. Therefore, our conceptual framework and empirical findings can best serve as a starting point for understanding and improving the creativity of frontier-size models today, as we begin to bridge the gap between human and machine intelligence.

大模型创造力生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。