arXiv:2605.10574cs.AI2026-05被引 1

发现大模型科学创造力呈现锯齿状分布,可被利用提升创新性能

LLM Jaggedness Unlocks Scientific Creativity

论文配图:LLM Jaggedness Unlocks Scientific Creativity
图 1 · 摘自论文原文
  • 构建科学创意评测基准SciAidanBench,量化模型生成独特科学想法的能力
  • 19个模型在不同任务、提示和领域间表现不均,展现显著锯齿状能力分布
  • 通过推理时计算、知识融合与头脑风暴,集成多模型实现超越单体的创造力

随着人工智能发展,模型能力并非均匀提升,而是呈现锯齿状动态,各任务、领域及规模下进展不一。本文聚焦科学创意生成,提出SciAidanBench基准,评估大语言模型(LLMs)在开放性科学问题上的创意潜力——要求模型生成尽可能多且一致的原创想法,以总有效响应数作为创造力代理指标。对8家厂商共19个基础模型(含30个变体,含推理版本)的评估显示:跨模型层面,通用创造力提升不必然带来科学创造力提升,能力分布各异;在提示层面,强模型表现波动大,部分问题爆发式创意,部分则表现平庸;在领域层面,模型在不同科学子领域中能力分布不均,体现内部能力碎片化。最后,我们证明这种锯齿性可被利用:通过推理时计算、知识池化与头脑风暴机制,有效组合模型,构建出超越任一单模型的元模型集成系统。结果表明,锯齿性不是缺陷,而是可被理解与利用的结构特征,能增强基于LLM的科学创造力。

原文摘要 · Abstract (English)

As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities growing unevenly across tasks, domains, and model scales. In this work, we examine this dynamic jaggedness through the lens of scientific idea generation. We introduce SciAidanBench, a benchmark of open-ended scientific questions designed to measure the scientific creativity of large language models (LLMs). Given a scientific question, models are asked to generate as many unique and coherent ideas as possible, with the total number of valid responses serving as a proxy for creative potential. Evaluating 19 base models across 8 providers (30 total variants including reasoning versions), we find that jaggedness manifests both across models and within models. First, in a cross-task comparison between general and scientific creativity, improvements in general creativity do not translate uniformly to scientific creativity, revealing divergent capability profiles across models. Second, at the prompt level, stronger models do not improve uniformly; instead, they exhibit high variability, with bursts of creativity on some questions and limited performance on others. Third, at the domain level, individual models display uneven strengths across scientific subfields, reflecting fragmented internal capability profiles. Finally, we show that this jaggedness can be harnessed. We explore mechanisms of inference-time compute, knowledge pooling, and brainstorming to combine models effectively and construct meta-model ensembles that outperform any single model. Our results position jaggedness not as a limitation, but as a resource, a structural feature of AI progress that, when understood and leveraged, can amplify LLM-driven scientific creativity.

科学创造大模型评测锯齿性模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。