提出新测试方法,更准确评估大模型的科学创意能力。
Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers

- 设计新测试DRAT,融合发散与收敛思维
- 现有测试无法可靠预测科学创意表现
- 适用于研究模型创造力的科研人员
衡量大语言模型(LLMs)的创造力对改进生成方法和理解该能力至关重要。近年来,常通过人类创造力测试来评估模型,但这些测试作为机器创造力指标的有效性尚未确立,且对人类创造力的预测效力也有限。为此,我们首次开展大规模系统研究,评估人类创造力测试在预测模型创造性写作、发散思维和科学构想三方面表现的效果。结果发现,发散联想任务(DAT)和条件发散联想任务(Conditional DAT)分别是创作和发散思维的最佳预测工具,但不同任务效果差异显著,无单一测试能全面覆盖所有维度。尤其值得注意的是,现有测试均无法可靠预测科学构想能力。为此,我们提出发散远距离联想测试(DRAT),一种基于词汇空间的新型测试,可同时评估收敛与发散思维。DRAT是首个能显著预测科学构想能力的模型创造力测试,且在多种设计选择下表现稳健。更重要的是,其性能无法通过DAT与远距离联想测试(RAT)线性组合恢复,说明将两类思维整合于同一测试中对科学创意预测至关重要。
原文摘要 · Abstract (English)
Measuring the creativity of large language models (LLMs) is essential for designing methods that can improve creativity and for enhancing our scientific understanding of this ability. To accomplish this, it has become common in recent years to administer tests of human creativity to LLMs. Although these tests provide a convenient and fully automated way to score "creativity," their validity as measures of machine creativity has not been established, and these tests already have limited validity as predictors of human creativity. To address this problem, we conduct the first large-scale, systematic study assessing the effectiveness of human creativity tests for predicting the creative achievement of LLMs across three target constructs: creative writing, divergent thinking, and scientific ideation. We find that the Divergent Association Task (DAT) and the Conditional DAT are the best predictors of creative writing and divergent thinking, respectively, but that test effectiveness varies significantly by construct, and no single test predicts all constructs well. Moreover, contrary to popular belief, no existing test reliably predicts scientific ideation ability. Motivated by this problem, we introduce the Divergent Remote Association Test (DRAT), a vocabulary-space test that assesses both convergent and divergent thinking in a single instrument. The DRAT is the first and only creativity test for LLMs that is a significant predictor of scientific ideation ability, demonstrating robustness across major design choices. Furthermore, the performance gain of the DRAT is not recoverable from any linear combination of the Divergent Association Task and the Remote Associates Test, indicating that assessing divergent and convergent thinking in the same test is essential to reliably predicting scientific ideation ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。