arXiv:2606.00875cs.CL2026-06

提出IDEAFix框架,系统评估大模型在创意生成中的发散思维能力。

IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs

论文配图:IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs
图 1 · 摘自论文原文
  • 设计可控任务与提示策略,研究结构化引导对创意生成的影响。
  • 发现提示策略可提升方案原创性,但模型输出仍高度同质化。
  • 适合关注大模型创造力评估与提示工程的研究者使用。

大语言模型在创造性问题解决和创意生成任务中应用日益广泛,但其创造力仍存争议:部分研究称其表现优于人类,也有研究指出存在思维固化与输出同质化等结构性缺陷。现有评估方法或局限于非情境化任务,或涵盖多因素混杂的宽泛场景,难以分离任务设计、提示策略与评估方式的影响。尤其缺乏对结构化提示如何塑造创意生成的深入研究。为此,本文提出IDEAFix——一个用于分析开放式创意生成中发散思维的评估框架。通过在控制变化的设计情景、任务属性及去固化提示策略下,引导模型生成多个原创解法,实现对结构化引导作用的系统分析。结果表明,任务设计与属性选择显著影响模型表现,简单提示策略可有效提升方案原创性。然而,不同模型间仍普遍存在输出同质化现象,证实其生成多样性存在固有局限。总体而言,IDEAFix提供了一个可控、可扩展的框架,有助于揭示大模型创造力的作用机制。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for tasks involving creative problem solving and idea generation. However, there is a lack of consensus concerning their creative capabilities: some studies report superior performances compared to humans, while others highlight structural limitations such as fixation and the homogenization of outputs. Existing evaluation approaches either rely on narrow, decontextualized tasks that do not capture goal-oriented generation or on broader settings that confound multiple aspects of the creative process, making it difficult to isolate the effects of task formulation, prompting, and evaluation design. Significantly, the role of structured prompting strategies in shaping idea generation remains underexplored. Therefore, we introduce IDEAFix, an evaluation framework for analyzing divergent thinking in open-ended idea generation tasks. We prompt models to generate multiple original solutions to controlled variations of short design scenarios, task attributes, and defixation prompting strategies. This design enables systematic analysis of how structured guidance influences LLMs' idea generation. Our results show that both task formulation and attribute selection significantly affect models' performance, and that simple prompting strategies can boost the originality of solutions. However, we also observe persistent output homogenization across models, confirming inherent limits in their ability to generate diverse solutions. Overall, IDEAFix provides a controlled, extensible framework for studying the mechanisms underlying LLMs' creativity.

大模型创意生成提示工程评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。