arXiv:2607.01233cs.CLcs.AI2026-07被引 4

对比大模型与人类研究思路,发现模型创意更集中于拼接融合。

Measuring the Gap Between Human and LLM Research Ideas

论文配图:Measuring the Gap Between Human and LLM Research Ideas
图 1 · 摘自论文原文
  • 从高质量论文反推灵感来源,构建评测框架
  • 大模型想法多集中在拼接式创新,人类则更广泛分布
  • 揭示模型创意在方向上系统性偏离人类研究偏好

大模型被越来越多用于生成研究想法,但现有评估多依赖新颖性、可行性或专家偏好。本文提出新视角:当前大模型生成的想法与人类研究思路有多远?为此,我们构建了基于高质量人类论文的规模化评测框架。对每篇论文,我们逆向推导其核心思想可能受哪些前期工作启发,并让大模型基于这些参考文献标题和摘要生成新想法。引入双轴研究品味分类法,以机会模式和研究范式刻画每个想法,量化人类与大模型想法的差异。结果显示,不同大模型生成的想法在分布上存在一致差距:大模型想法过度集中于桥梁型机会和综合方法,而人类论文的参考文献分布更广泛地覆盖问题定义方式与贡献构建路径。这表明强大多模型虽能生成多样合理想法,但其创意范围仍比人类狭窄,且系统性偏离人类研究品味。

原文摘要 · Abstract (English)

LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from human researchers? To characterize this gap, we build a large-scale evaluation framework for ideation from high-quality human research papers. For each paper, we reverse-engineer a small set of closely related prior works that likely inspired its core idea. LLMs are then prompted to generate a new idea from the set of paper titles and summaries. We introduce a two-axis research-taste taxonomy to profile each idea by its opportunity pattern and research paradigm, and use it to quantify the divergence between human and LLM ideas. Across idea sets generated by different LLMs, we observe a consistent distributional gap: LLM ideas are disproportionately concentrated around bridge-like opportunities and synthesis methods, whereas the human paper reference distribution spreads more broadly across ways of framing gaps and constructing contributions. This result suggests that strong LLMs can produce a range of reasonable ideas, but that range remains narrower than, and systematically shifted relative to, human research taste.

大模型创意研究思维评测框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。