arXiv:2510.12110cs.CLcs.AI2025-10EMNLP被引 9

用平行联想链评估大模型创造力,高效且可靠。

Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models

  • 让模型生成平行联想链,通过结构化输出衡量创意水平。
  • 与人工评分相关性达0.739,显著优于传统方法。
  • 适合关注模型创意能力评估的研究者和开发者。

大语言模型(LLM)的创造力评估是重要研究方向,但数据污染和高昂的人工评估成本常成阻碍。受人类创造力评估启发,我们提出PACE方法,要求模型生成平行联想链以评估其创造力。该方法有效降低数据污染风险,具备高效、简洁的特点。实验显示,PACE在多个专有及开源模型上与Chatbot Arena创意写作排名呈现强相关性(斯皮尔曼等级相关系数ρ=0.739,p<0.001)。对比分析发现,高性能大模型的联想创造力可达到普通人类水平,但专业人类仍持续领先。语言学分析进一步表明,人类与模型均呈现联想抽象性上升趋势,但人类表现出更丰富的联想模式多样性。

原文摘要 · Abstract (English)

The evaluation of LLMs' creativity represents a crucial research domain, though challenges such as data contamination and costly human assessments often impede progress. Drawing inspiration from human creativity assessment, we propose PACE, asking LLMs to generate Parallel Association Chains to Evaluate their creativity. PACE minimizes the risk of data contamination and offers a straightforward, highly efficient evaluation, as evidenced by its strong correlation with Chatbot Arena Creative Writing rankings (Spearman's $ρ= 0.739$, $p < 0.001$) across various proprietary and open-source models. A comparative analysis of associative creativity between LLMs and humans reveals that while high-performing LLMs achieve scores comparable to average human performance, professional humans consistently outperform LLMs. Furthermore, linguistic analysis reveals that both humans and LLMs exhibit a trend of decreasing concreteness in their associations, and humans demonstrating a greater diversity of associative patterns.

大模型评估创造力联想链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。