用大模型预测AI研究想法成败,准确率超人类专家。
Predicting Empirical AI Research Outcomes with Language Models
- 结合微调GPT-4.1与论文检索,自动判断两个研究想法谁更可能成功。
- 在NLP领域准确率达64.4%,全集测试达77%,显著优于人类专家。
- 可评估未发表想法,适合用于优化AI创意生成模型的奖励机制。
许多看似有前景的AI研究想法最终未能实现预期效果,但验证这些想法需大量人力与算力。因此,预测其成功概率对加速实证性AI研究至关重要,这一能力通常需长期经验积累。本文构建首个该任务基准,对比大模型与人类专家表现。具体而言,给定两个研究想法(如两种越狱方法),目标是预测哪个在一组基准上表现更好。我们从会议论文中爬取1,585对经人工验证的想法配对(发布于基线模型截断日期后)用于测试,6,000对用于训练。开发系统结合微调GPT-4.1与论文检索代理,并招募25名人类专家参与对比。在NLP领域,系统准确率达64.4%,显著优于人类专家的48.9%;全测试集准确率达77%。而现成前沿大模型如o3,即使加入相同检索增强,表现仍接近随机猜测。通过大量人工撰写与大模型设计的鲁棒性测试,验证系统未依赖复杂度等表面特征。最后,系统对未发表新想法(包括由AI创意代理生成者)进行评估,准确率达63.6%,证明其作为改进创意生成模型奖励模型的潜力。整体结果揭示大模型加速实证研究的新方向。
原文摘要 · Abstract (English)
Many promising-looking ideas in AI research fail to deliver, but their validation takes substantial human labor and compute. Predicting an idea's chance of success is thus crucial for accelerating empirical AI research, a skill that even expert researchers can only acquire through substantial experience. We build the first benchmark for this task and compare LMs with human experts. Concretely, given two research ideas (e.g., two jailbreaking methods), we aim to predict which will perform better on a set of benchmarks. We scrape ideas and experimental results from conference papers, yielding 1,585 human-verified idea pairs published after our base model's cut-off date for testing, and 6,000 pairs for training. We then develop a system that combines a fine-tuned GPT-4.1 with a paper retrieval agent, and we recruit 25 human experts to compare with. In the NLP domain, our system beats human experts by a large margin (64.4% v.s. 48.9%). On the full test set, our system achieves 77% accuracy, while off-the-shelf frontier LMs like o3 perform no better than random guessing, even with the same retrieval augmentation. We verify that our system does not exploit superficial features like idea complexity through extensive human-written and LM-designed robustness tests. Finally, we evaluate our system on unpublished novel ideas, including ideas generated by an AI ideation agent. Our system achieves 63.6% accuracy, demonstrating its potential as a reward model for improving idea generation models. Altogether, our results outline a promising new direction for LMs to accelerate empirical AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。