让大模型自主学习科研技能,突破工具依赖限制
CASCADE: Cumulative Agentic Skill Creation through Autonomous Development and Evolution
- 通过搜索、代码提取和记忆利用实现持续学习
- 在材料化学任务中达93.3%成功率,无进化机制仅35.4%
- 适合科研自动化与跨代理技能共享场景
当前大型语言模型(LLM)代理依赖预定义工具或早期工具生成,限制其在复杂科学任务中的适应性和可扩展性。我们提出CASCADE,一种自演化智能体框架,标志着从‘LLM+工具使用’向‘LLM+技能获取’的早期演进。CASCADE使代理能够掌握复杂外部工具并编码知识,依赖两种元技能:通过网络搜索、代码提取和记忆利用实现持续学习;通过自我反思、知识图谱探索等实现自我审视。我们在SciSkillBench(包含116个材料科学与化学研究任务的基准)上评估,使用GPT-5时成功率达93.3%,而无演化机制下仅为35.4%。我们进一步展示了在计算分析、自主实验和复现已发表论文中的实际应用。结合人机协作与记忆整合,CASCADE可累积可执行技能并在代理与科学家间共享,推动可扩展的AI辅助科学研究。
原文摘要 · Abstract (English)
Large language model (LLM) agents currently depend on predefined tools or early-stage tool generation, limiting their adaptability and scalability to complex scientific tasks. We introduce CASCADE, a self-evolving agentic framework representing an early instantiation of the transition from "LLM + tool use" to "LLM + skill acquisition". CASCADE enables agents to master complex external tools and codify knowledge through two meta-skills: continuous learning via web search, code extraction, and memory utilization; self-reflection via introspection, knowledge graph exploration, and others. We evaluate CASCADE on SciSkillBench, a benchmark of 116 materials science and chemistry research tasks. CASCADE achieves a 93.3% success rate using GPT-5, compared to 35.4% without evolution mechanisms. We further demonstrate real-world applications in computational analysis, autonomous laboratory experiments, and selective reproduction of published papers. Along with human-agent collaboration and memory consolidation, CASCADE accumulates executable skills that can be shared across agents and scientists, moving toward scalable AI-assisted scientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。