arXiv:2608.07243cs.AIcs.CL2026-08中稿 · ICCC'26

用迭代生成与评估提升大模型菜谱创意,发现评价器设计比迭代次数更重要。

Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

论文配图:Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models
图 1 · 摘自论文原文
  • 通过多轮生成-筛选循环优化菜谱创意,模拟人类创作过程。
  • 小规模评分模型在多数创意维度上表现更优,迭代次数增效不明显。
  • 适合关注创意生成机制与评价体系的AI研究员或内容创作者。

生成模型通常依赖单一输出进行评估,而人类创造力往往通过反复生成、评价与改进产生。本初步研究将FunSearch方法适配至2024年Pillsbury Bake-Off菜谱生成任务,采用基于TTCT的LLM评估框架,考察迭代次数、生成温度及环内选择评分模型规模的影响。两个实验结果表明,迭代生成-选择机制可产出创意得分接近人类基准的菜谱,但单纯增加迭代次数无法进一步提升创意。最关键的是,较小的环内评分模型在多数TTCT维度上显著优于大模型,而温度仅对原创性有有限影响。结果表明,评价器设计是主观创意搜索中的首要设计变量。

原文摘要 · Abstract (English)

Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting FunSearch to recipe generation for the 2024 Pillsbury Bake-Off and evaluating outputs against human benchmarks using TTCT-based LLM evaluation. Across two experiments, we test iteration count, generator temperature, and in-loop selection-scorer model size. Results show that iterative generation-selection can produce recipes with creativity scores comparable to human benchmarks, but additional iterations alone do not improve creativity. The in-loop evaluator matters most: a smaller selection scorer yields significantly higher scores across most TTCT dimensions, while temperature has limited effects except for originality. These findings suggest that evaluator design is a first-order design variable in subjective creative search.

创意生成大模型评估迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。