用引导式生成提升科学创意的效率与质量
Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents

- 以参考创意为引导,定向生成科研推理路径
- 合成数据样本效率提升10倍以上,质量高11.91%
- 适合想高效构建科研智能体的研究者
创意在科学发现中至关重要。近期大模型,尤其是AI科学家系统,在自动化创意生成方面展现出巨大潜力。然而,现有方法多依赖预设的智能体工作流,限制了在浩瀚文献与复杂推理空间中的灵活性。训练自主智能体大模型成为新方向,具备灵活推理和自主调用工具的能力。但此前的智能体数据生成方法在科学创意场景中存在数据生成成本过高问题。为此,我们提出Agentic-Ideation框架,包含自动化轨迹生成流水线与专用于科学创意的智能体大模型。首先定义涵盖三种外部工具和三种认知工具的综合工具空间;随后引入基于参考创意的引导式数据生成策略,通过引导多智能体系统高效重构逻辑推理与工具调用路径,将无序试错转化为有向轨迹生成。最后,基于这些合成轨迹训练模型,并对工具执行结果施加掩码策略,使模型专注决策逻辑而非外部反馈。实验表明,该方法在整体质量上优于最先进工作流基线11.91%,且高质量数据合成样本效率提升超10倍。
原文摘要 · Abstract (English)
Ideation plays a pivotal role in scientific discovery. Recent LLM, especially AI Scientist systems, show promising potential for automated ideation. However, existing approaches predominantly rely on pre-defined agentic workflows. This constraint severely limits the flexibility required to navigate the vast search space of scientific literature and the complex action space of research reasoning. Recently, training Agentic LLMs has emerged as a promising direction, offering flexible reasoning frameworks and the capability for autonomous tool utilization. However, there remains a non-trivial challenge: applying previous agentic data synthesis methods to scientific ideation suffers from prohibitively high data synthesis cost. To bridge this gap, we propose Agentic-Ideation, a novel framework comprising an automated trajectory synthesis pipeline and a specialized agentic LLM trained for scientific ideation. Specifically, we first define a comprehensive tool space incorporating three external tools and three cognitive tools. Then we introduce an Oracle-Guided Data Synthesis strategy. By leveraging a reference idea as oracle guidance, this approach steers the multi-agent system to efficiently reconstruct the logical reasoning and tool invocation paths, transforming aimless trial-and-error into directed trajectory generation. Finally, we train the agent on these synthesized trajectories, employing a masking strategy on tool execution results. This ensures the model focuses on decision-making logic without interference from external feedback. Experimental results demonstrate that our method outperforms the SOTA workflow-based baseline by \textbf{11.91\%} in overall quality. Furthermore, our approach improves the sample efficiency of high-quality data synthesis by \textbf{over 10$\times$}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。