arXiv:2603.11863cs.AIcs.CL2026-03ACL被引 9

构建可自进化挑战的评测基准,量化机器代码创造力

CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges

  • 基于逆向工程与自对弈的自动化流水线生成创意任务
  • 大模型规模提升显著增强组合式创造力,但探索能力趋于饱和
  • 提出EvoRePE策略,在推理阶段持续提升创造力,适合生成类任务

高质量预训练数据的饱和促使研究转向能够持续生成新成果的演化系统,如AlphaEvolve。然而,这类系统的进展受限于缺乏严格的量化评估。为此,我们提出CreativeBench,一个基于经典认知框架的代码生成创造力评测基准。该基准包含两个子集:CreativeBench-Combo(组合型)和CreativeBench-Explore(探索型),通过逆向工程与自对弈的自动化流程构建挑战。利用可执行代码,以质量与新颖性的乘积作为统一指标,客观区分创造力与幻觉。对主流模型的分析显示:(1) 模型规模显著提升组合创造力,但探索能力收益递减;(2) 更大模型呈现“缩放趋同”,更准确但更保守;(3) 推理能力主要促进受约束的探索而非组合。最后,我们提出EvoRePE,一种即插即用的推理时引导策略,内化演化搜索模式,持续增强机器创造力。

原文摘要 · Abstract (English)

The saturation of high-quality pre-training data has shifted research focus toward evolutionary systems capable of continuously generating novel artifacts, leading to the success of AlphaEvolve. However, the progress of such systems is hindered by the lack of rigorous, quantitative evaluation. To tackle this challenge, we introduce CreativeBench, a benchmark for evaluating machine creativity in code generation, grounded in a classical cognitive framework. Comprising two subsets -- CreativeBench-Combo and CreativeBench-Explore -- the benchmark targets combinatorial and exploratory creativity through an automated pipeline utilizing reverse engineering and self-play. By leveraging executable code, CreativeBench objectively distinguishes creativity from hallucination via a unified metric defined as the product of quality and novelty. Our analysis of state-of-the-art models reveals distinct behaviors: (1) scaling significantly improves combinatorial creativity but yields diminishing returns for exploration; (2) larger models exhibit ``convergence-by-scaling,'' becoming more correct but less divergent; and (3) reasoning capabilities primarily benefit constrained exploration rather than combination. Finally, we propose EvoRePE, a plug-and-play inference-time steering strategy that internalizes evolutionary search patterns to consistently enhance machine creativity.

创造力评测代码生成演化系统推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。