用多阶段AI流程生成更原创的研究计划,效果优于单一提示。
Evaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines
- 设计多步智能体流程,通过分解与长上下文增强创意
- 分解和长上下文流程新颖性得分4.17/5,远高于反思型的2.33/5
- 适合科研助手、创新策划者及需要高质量研究思路的人
大型语言模型融入科研生态引发对AI生成研究创意原创性的根本疑问。近期研究指出,单步提示存在‘智能抄袭’问题,即模型通过术语替换复现已有想法。本文探究多步智能体工作流——包含迭代推理、进化搜索与递归分解——能否生成更具新颖性与可行性研究计划。我们评测了五种推理架构:基于反思的迭代优化、Sakana AI v2进化算法、Google Co-Scientist多智能体框架、GPT Deep Research(GPT-5.1)递归分解,以及Gemini~3 Pro多模态长上下文流水线。每种方法在三十个提案上评估新颖性、可行性和影响力。结果表明,基于分解与长上下文的工作流平均新颖性达4.17/5,而反思型方法仅2.33/5。不同研究领域表现各异,高性能工作流在保持可行性的同时未牺牲创造力。研究支持精心设计的多阶段智能体流程可推动AI辅助科研构思。
原文摘要 · Abstract (English)
The integration of Large Language Models (LLMs) into the scientific ecosystem raises fundamental questions about the creativity and originality of AI-generated research. Recent work has identified ``smart plagiarism'' as a concern in single-step prompting approaches, where models reproduce existing ideas with terminological shifts. This paper investigates whether agentic workflows -- multi-step systems employing iterative reasoning, evolutionary search, and recursive decomposition -- can generate more novel and feasible research plans. We benchmark five reasoning architectures: Reflection-based iterative refinement, Sakana AI v2 evolutionary algorithms, Google Co-Scientist multi-agent framework, GPT Deep Research (GPT-5.1) recursive decomposition, and Gemini~3 Pro multimodal long-context pipeline. Using evaluations from thirty proposals each on novelty, feasibility, and impact, we find that decomposition-based and long-context workflows achieve mean novelty of 4.17/5, while reflection-based approaches score significantly lower (2.33/5). Results reveal varied performance across research domains, with high-performing workflows maintaining feasibility without sacrificing creativity. These findings support the view that carefully designed multi-stage agentic workflows can advance AI-assisted research ideation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。