arXiv:2602.01550cs.AI2026-02

自进化科研智能体框架,能持续学习并高效调度跨学科工具。

S1-NexusAgent: a Self-Evolving Agent Framework for Multidisciplinary Scientific Research

  • 分层规划+代码执行,双循环架构实现复杂科研流程稳定运行。
  • 在生物、化学、材料三大学科基准上达当前最优性能。
  • 自动提炼科研经验形成可复用技能,适合长期科研项目使用。

现代科学研究依赖大规模数据、复杂工作流和专用工具,现有大模型与工具型智能体因长程规划能力弱、目标保持不稳及持续学习不足而难以应对。为此,本文提出S1-NexusAgent——一种面向多学科科研的自进化智能体框架。该框架采用分层规划-代码执行范式,通过双循环架构将全局科研规划与子任务工具执行解耦,实现复杂研究流程的稳定建模。系统原生支持模型上下文协议(MCP),集成超千个跨学科科学工具,利用意图感知的动态工具检索与热插拔机制,高效编排异构研究工具。为应对科研场景中的长上下文与大规模数据挑战,引入基于对象引用的稀疏上下文管理,实现子任务上下文隔离与中间结果压缩。在此基础上,批判者智能体自动评估完整执行轨迹,将高质量科研路径提炼为可复用的科学技能,形成持续自我演化的闭环,对可持续长周期科研具有重要意义。在包含生物(biomini-eval)、化学(ChemBench)和材料科学(MatSciBench)的权威基准测试中,S1-NexusAgent在长程规划与复杂专用工具编排任务上均达到当前最优表现,验证了其有效性与泛化能力。

原文摘要 · Abstract (English)

Modern scientific research relies on large-scale data, complex workflows, and specialized tools, which existing LLMs and tool-based agents struggle to handle due to limitations in long-horizon planning, robust goal maintenance, and continual learning from execution. To address these issues, in this work, we propose S1-NexusAgent, a self-evolving agent framework designed for multidisciplinary scientific research. S1-NexusAgent adopts a hierarchical Plan-and-CodeAct execution paradigm, decoupling global scientific planning from subtask-level tool execution through a dual-loop architecture, thereby enabling stable modeling of complex research workflows. The system natively supports the Model Context Protocol (MCP), integrates up to thousands of cross-disciplinary scientific tools, and achieves efficient orchestration of heterogeneous research tools via intention-aware dynamic tool retrieval and hot-plug mechanisms. To address long-context and large-scale data challenges in scientific settings, S1-NexusAgent introduces object-reference-based sparse context management, which enables sub-task context isolation and intermediate result compression. Building on this, a Critic Agent automatically evaluates complete execution trajectories and distills high-quality research paths into reusable Scientific Skills, forming a closed loop for continuous self-evolution, which is valuable for sustainable and long-horizon scientific research. Experiments on authoritative scientific benchmarks involving long-horizon planning and complex specialized tool orchestration, including biomini-eval (biology), ChemBench (chemistry), and MatSciBench (material science), demonstrate that S1-NexusAgent achieves state-of-the-art performance, validating its effectiveness and generalization capability in complex scientific tasks.

智能体科研自动化自进化多学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。