用AI把科研点子变成完整计划,测试大模型的科研规划能力。
Idea2Plan: Exploring AI-Powered Research Planning
- 构建Idea2Plan任务与基准,评估大模型从想法到计划的转化能力。
- GPT-5表现最佳但仍有提升空间,表明当前模型规划能力有限。
- 适合研究AI科研助手、自主研究代理的开发者参考。
大型语言模型(LLMs)在加速科学发现方面展现出巨大潜力,可辅助数据分析、假设生成及创新方法探索。本文研究LLMs如何将抽象科研想法转化为结构化研究计划。有效科研规划不仅助力科学家推进研究,也是发展自主科研代理的关键能力。然而,该领域尚缺乏对LLMs科研规划能力的系统理解。为此,我们提出Idea2Plan任务与Idea2Plan Bench基准,基于ICML 2025和Nature Mental Health论文构建,这些论文均发布于主要LLM训练截止日期之后。每个基准实例包含一个研究想法与评分标准,涵盖有效计划的核心要素。我们还提出Idea2Plan JudgeEval,用于评估基于LLM的评判者相对于专家标注的可靠性。实验结果表明,GPT-5在基准上表现最强,但仍存在显著提升空间。本研究为LLMs科研规划能力提供了新洞见,并为未来进展奠定基础。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated significant potential to accelerate scientific discovery as valuable tools for analyzing data, generating hypotheses, and supporting innovative approaches in various scientific fields. In this work, we investigate how LLMs can handle the transition from conceptual research ideas to well-structured research plans. Effective research planning not only supports scientists in advancing their research but also represents a crucial capability for the development of autonomous research agents. Despite its importance, the field lacks a systematic understanding of LLMs' research planning capability. To rigorously measure this capability, we introduce the Idea2Plan task and Idea2Plan Bench, a set of benchmarks built from ICML 2025 and Nature Mental Health papers released after major LLM training cutoffs. Each benchmark instance includes a research idea and a grading rubric capturing the key components of valid plans. We further propose Idea2Plan JudgeEval, a complementary benchmark to assess the reliability of LLM-based judges against expert annotations. Experimental results show that GPT-5 achieves the strongest performance on the benchmark, though substantial headroom remains for improvement. Our study provides new insights into LLMs' capability for research planning and lays the groundwork for future progress.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。