LLM生成的研究创意比专家更新颖,但可行性略低。
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- 让100多位专家与LLM对比生成研究创意,控制变量评估。
- LLM创意新颖性显著高于人类(p<0.05),可行性稍弱。
- 揭示了LLM自我评估失效与创意多样性不足的问题。
大语言模型(LLMs)的进展引发了对其加速科学发现潜力的期待,已有研究提出可自主生成并验证新想法的研究代理。然而,尚无实证证明LLM能完成研究的第一步——生成专家级的新颖创意,更不用说完成整个研究流程。为此,我们设计了一项实验,控制混杂因素,首次实现专家自然语言处理(NLP)研究人员与一个LLM创意生成代理的直接比较。通过招募超过100位NLP研究人员撰写新颖创意,并对人类与LLM创意进行盲评,我们得出当前关于研究创意生成的首个统计显著结论:LLM生成的创意在新颖性上优于人类专家创意(p < 0.05),但在可行性上略逊一筹。深入分析代理基线后,我们识别出构建与评估研究代理的关键问题,包括LLM自我评估失败及生成多样性不足。此外,我们指出专家对新颖性的判断本身具有挑战性,并提出一种端到端研究设计,即招募研究人员将这些创意付诸实践,以检验新颖性与可行性判断是否导致实际研究结果的差异。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have sparked optimism about their potential to accelerate scientific discovery, with a growing number of works proposing research agents that autonomously generate and validate new ideas. Despite this, no evaluations have shown that LLM systems can take the very first step of producing novel, expert-level ideas, let alone perform the entire research process. We address this by establishing an experimental design that evaluates research idea generation while controlling for confounders and performs the first head-to-head comparison between expert NLP researchers and an LLM ideation agent. By recruiting over 100 NLP researchers to write novel ideas and blind reviews of both LLM and human ideas, we obtain the first statistically significant conclusion on current LLM capabilities for research ideation: we find LLM-generated ideas are judged as more novel (p < 0.05) than human expert ideas while being judged slightly weaker on feasibility. Studying our agent baselines closely, we identify open problems in building and evaluating research agents, including failures of LLM self-evaluation and their lack of diversity in generation. Finally, we acknowledge that human judgements of novelty can be difficult, even by experts, and propose an end-to-end study design which recruits researchers to execute these ideas into full projects, enabling us to study whether these novelty and feasibility judgements result in meaningful differences in research outcome.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。