构建跨领域大模型创意评估框架,揭示模型在不同创意维度上的表现差异。
CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity
- 整合三大领域八项任务,从质量、新颖性、多样性三维度评估创意。
- 前沿大模型在写作与逻辑推理上领先开源模型15%,但发散思维能力无优势。
- 各创意维度间相关性弱,需跨领域框架才能全面评估模型创造力。
创意常被视为人类智能的标志。尽管大语言模型(LLMs)被广泛认为能生成创意文本,但目前仍缺乏跨领域且可扩展的评估框架来衡量其在多样化场景下的创意表现。现有方法或严重依赖人工评价,影响速度与可扩展性;或在不同领域和创意定义间碎片化。为此,我们提出CreativityPrism,一个整合了发散思维、创意写作与逻辑推理三个领域共八项任务的评估与分析框架,构建以质量、新颖性与多样性为核心的创意分类体系。该框架具备可扩展性,采用经人工标注验证的自动评价判别器。我们在17个当前最先进的(SoTA)LLM上测试该框架,发现前沿大模型在创意写作与逻辑推理任务中比本地部署的开源模型领先15%(即0.10分),但在发散思维这一较少被现有训练范式覆盖的领域并无显著优势。分析还显示,某一创意维度或领域的高表现极少能推广至其他维度或领域,尤其新颖性指标与其他指标常呈弱相关甚至负相关。该结果证实,像CreativityPrism这样的跨领域、多维度框架对实现对大模型创意能力的有意义评估至关重要。
原文摘要 · Abstract (English)
Creativity is often seen as a hallmark of human intelligence. While large language models(LLMs) are increasingly perceived as generating creative text, there is still no cross-domain and scalable framework to evaluate their creativity across diverse scenarios. Existing methods of LLM creativity evaluation either heavily rely on humans, limiting speed and scalability, or are fragmented across different domains and different definitions of creativity. To address this gap, we propose CreativityPrism, an evaluation and analysis framework that consolidates eight tasks from three domains: divergent thinking, creative writing, and logical reasoning, into a taxonomy of creativity that emphasizes three dimensions: quality, novelty, and diversity of LLM generations. The framework is designed to be scalable with reliable automatic evaluation judges that have been validated against human annotations. We evaluate 17 state-of-the-art (SoTA) LLMs on CreativityPrism and find that while frontier-scale LLMs dominate creative writing and logical reasoning tasks by a .10 (or 15%) lead over locally-deployable open models, they offer no significant advantage in divergent thinking, a domain much less explored in existing post-training regimes. Our analysis also shows that high performance in one creative dimension or domain rarely generalizes to others; specifically, novelty metrics often show weak or negative correlations with other metrics. This fragmentation confirms that a cross-domain, multi-dimensional framework like CreativityPrism is essential for any meaningful assessment of LLM creativity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。