用关键词测试大模型科学创意,发现创造力与通用能力无关。
Evaluating LLMs' Divergent Thinking Capabilities for Scientific Idea Generation with Minimal Context
- 仅用一个关键词评估模型的发散性思维能力。
- 40多个模型在1180个关键词上测试,发现创意表现与通用能力不相关。
- 提示:适合关注AI辅助科研、创意生成的研究者。
尽管大语言模型在文献分析和实验设计等科学任务中表现优异(如准确提取论文关键发现或生成连贯实验流程),现有评估基准主要依赖丰富上下文输入。我们提出LiveIdeaBench,一个全面的基准,通过单关键词提示评估模型的科学创意生成能力,基于吉尔福德创造力理论,从原创性、可行性、流畅性、灵活性和清晰度五个维度评估。在22个科学领域共1180个关键词上对超过40个领先模型进行大规模实验,结果表明:本基准测得的科学创意能力无法被通用智能指标预测。例如,QwQ-32B-preview的创意表现可媲美顶级模型claude-3.7-sonnet:thinking,尽管其通用智能得分存在显著差距。研究提示需针对科学创意生成开发专用评估框架,并可能需要不同于通用问题求解的训练策略,以支持科研全周期的AI工具开发。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) demonstrate remarkable capabilities in scientific tasks such as literature analysis and experimental design (e.g., accurately extracting key findings from papers or generating coherent experimental procedures), existing evaluation benchmarks primarily assess performance using rich contextual inputs. We introduce LiveIdeaBench, a comprehensive benchmark evaluating LLMs' scientific idea generation by assessing divergent thinking capabilities using single-keyword prompts. Drawing from Guilford's creativity theory, our benchmark employs a dynamic panel of state-of-the-art LLMs to assess generated ideas across five key dimensions: originality, feasibility, fluency, flexibility, and clarity. Through extensive experimentation with over 40 leading models across 1,180 keywords spanning 22 scientific domains, we reveal that the scientific idea generation capabilities measured by our benchmark, are poorly predicted by standard metrics of general intelligence. Our results demonstrate that models like QwQ-32B-preview achieve creative performance comparable to top-tier models such as claude-3.7-sonnet:thinking, despite significant gaps in their general intelligence scores. These findings highlight the need for specialized evaluation benchmarks for scientific idea generation and suggest that enhancing these idea generation capabilities in LLMs may require different training strategies than those used for improving general problem-solving abilities, potentially enabling a wider range of AI tools tailored for different stages of the scientific process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。