AI生成的科研点子看似新颖,实则执行后效果不如人类想法。
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- 让43位专家分别实现人类和AI生成的研究点子
- 执行后AI点子评分大幅下降,多项指标落后于人类点子
- 揭示了当前AI创意在实际研究中的局限性
大型语言模型(LLMs)在加速科研流程方面展现出潜力,尤其在生成新研究思路方面。已有研究表明,在构思阶段,某些情境下由LLM生成的想法被认为比人类专家的想法更具新颖性。然而,一个好想法不仅需要看起来新颖,更应能在实际执行中取得更好成果。为此,我们开展了一项执行研究:招募43位专家级研究人员,随机分配执行由专家撰写或由LLM生成的研究想法,每位专家投入超过100小时进行实现,并撰写4页短论文记录实验过程。所有项目由不知情的NLP领域专家盲审。对比执行前后的评审分数,发现LLM生成想法在新颖性、吸引力、有效性及整体评分上显著低于人类想法(所有指标p < 0.05),弥补了构思阶段的差距。进一步统计显示,许多指标排名反转,人类想法反而优于AI生成想法。这一‘构想到执行’的差距揭示了当前LLM在生成真正有效的研究思路方面的局限,也凸显了仅靠构思评估研究想法的不足。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown promise in accelerating the scientific research pipeline. A key capability for this process is the ability to generate novel research ideas, and prior studies have found settings in which LLM-generated research ideas were judged as more novel than human-expert ideas. However, a good idea should not simply appear to be novel, it should also result in better research after being executed. To test whether AI-generated ideas lead to better research outcomes, we conduct an execution study by recruiting 43 expert researchers to execute randomly-assigned ideas, either written by experts or generated by an LLM. Each expert spent over 100 hours implementing the idea and wrote a 4-page short paper to document the experiments. All the executed projects are then reviewed blindly by expert NLP researchers. Comparing the review scores of the same ideas before and after execution, the scores of the LLM-generated ideas decrease significantly more than expert-written ideas on all evaluation metrics (novelty, excitement, effectiveness, and overall; p < 0.05), closing the gap between LLM and human ideas observed at the ideation stage. When comparing the aggregated review scores from the execution study, we even observe that for many metrics there is a flip in rankings where human ideas score higher than LLM ideas. This ideation-execution gap highlights the limitations of current LLMs in generating truly effective research ideas and the challenge of evaluating research ideas in the absence of execution outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。