让科研工具代理自动学习新工具并动态适应开放世界任务。
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

- 基于本体和持续记忆,动态发现与整合新科研工具。
- 在900个任务上达到当前最优性能,跨难度泛化能力强。
- 适合需要持续学习新工具的自动化科研场景。
大语言模型代理已广泛用于科学研宄中组织与调用专用计算工具。然而,其对预定义工具空间和静态语义的依赖限制了其在开放世界科研工作流中的应用,后者中工具需求、能力与边界会动态变化。为此,我们提出SciToolAgent-Evo,一种面向开放世界科学工具获取的本体感知自演化代理。该代理通过持续积累的技能、经验与本体化工具图,从对比轨迹中提炼可迁移知识;推理时采用基于LinUCB的强化学习门控机制,动态平衡探索与利用。一旦获取新工具,其科学本体将在线构建,实现无缝集成。此外,我们引入OpenSciToolBench基准,包含4个难度等级共900个真实任务。大量实验表明,SciToolAgent-Evo表现领先,验证了其鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on predefined tool spaces with static semantics limits their applicability to open-world scientific workflows, where tool requirements, capabilities, and boundaries evolve dynamically. To this end, we propose SciToolAgent-Evo, an ontology-aware self-evolving agent for open-world scientific tool acquisition. Driven by an evolving memory of skills, experiences, and an ontologized tool graph, it distills generalizable knowledge from contrastive trajectories during accumulation, whereas during inference, it formulates active requests and utilizes a LinUCB-based bandit gate to dynamically balance exploration and exploitation. Once a novel tool is acquired, its scientific ontology is completed online for seamless integration into the known graph. Moreover, we introduce OpenSciToolBench, a benchmark containing 900 realistic tasks across four difficulty levels. Extensive evaluations show that SciToolAgent-Evo achieves state-of-the-art performance, validating its robustness and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。