arXiv:2603.28986cs.AIcs.LG2026-03被引 1

Mimosa让科研多智能体系统自动演化,提升复杂任务成功率。

Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research

  • 通过动态工具发现与元协调器生成可迭代优化的科研工作流
  • 在ScienceAgentBench上达43.1%成功率,优于静态多智能体方案
  • 支持开源扩展与全程可审计,适合跨学科自动化科研

当前自主科学研究所依赖的大语言模型与智能体架构受限于固定流程与工具集,难以适应动态任务。我们提出Mimosa框架,通过模型上下文协议(MCP)实现动态工具发现,由元协调器生成工作流拓扑,利用代码生成智能体调用工具与科学库执行子任务,并以基于LLM的评判器评估结果,反馈驱动工作流持续优化。在ScienceAgentBench测试中,使用DeepSeek-V3.2的Mimosa达成43.1%的成功率,超越单智能体基线与静态多智能体配置。实验表明,不同模型对多智能体分解和迭代学习响应各异,说明工作流演化的收益取决于底层执行模型能力。Mimosa采用模块化、工具无关设计,支持快速扩展;完整执行日志与归档工作流保障可审计性,便于复现与审查。结合领域专家指导,该框架有望自动化跨学科计算型科研任务。项目已开源,旨在为社区共建自主科研提供开放基础。

原文摘要 · Abstract (English)

Current Autonomous Scientific Research (ASR) systems, despite leveraging large language models (LLMs) and agentic architectures, remain constrained by fixed workflows and toolsets that prevent adaptation to evolving tasks and environments. We introduce Mimosa, an evolving multi-agent framework that automatically synthesizes task-specific multi-agent workflows and iteratively refines them through experimental feedback. Mimosa leverages the Model Context Protocol (MCP) for dynamic tool discovery, generates workflow topologies via a meta-orchestrator, executes subtasks through code-generating agents that invoke available tools and scientific software libraries, and scores executions with an LLM-based judge whose feedback drives workflow refinement. On ScienceAgentBench, Mimosa achieves a success rate of 43.1% with DeepSeek-V3.2, surpassing both single-agent baselines and static multi-agent configurations. Our results further reveal that models respond heterogeneously to multi-agent decomposition and iterative learning, indicating that the benefits of workflow evolution depend on the capabilities of the underlying execution model. Beyond these benchmarks, Mimosa modular architecture and tool-agnostic design make it readily extensible, and its fully logged execution traces and archived workflows support auditability by preserving every analytical step for inspection and potential replication. Combined with domain-expert guidance, the framework has the potential to automate a broad range of computationally accessible scientific tasks across disciplines. Released as a fully open-source platform, Mimosa aims to provide an open foundation for community-driven ASR.

多智能体科研自动化自进化开源框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。