AI自主生成可复用的量子模拟工具,提升科学计算效率与精度。
El Agente Forjador: Task-Driven Agent Generation for Quantum Simulation

- 多智能体框架通过分析、生成、执行、评估四阶段自动构建代码工具。
- 在24个任务中,工具复用使弱模型准确率显著提升,API成本降低。
- 跨领域工具组合可解决混合任务,适用于需动态适配的新科学场景。
人工智能助力科学发现,大语言模型与智能体工作流加速了各类科学任务。然而现有系统依赖静态人工工具集,难以适应新领域和演进的库。我们提出El Agente Forjador,一个由通用编码智能体构成的多智能体框架,通过四阶段流程(工具分析、生成、任务执行、迭代评估)实现工具的自主锻造、验证与复用。在五个编码智能体配置下,对24个涵盖量子化学与量子动力学的任务进行评估,比较三种模式:按任务零样本生成工具、复用课程构建的工具集、以及直接问题求解作为基线。结果表明,工具生成与复用框架始终优于基线。同时,使用强智能体构建的工具集可降低弱智能体的API成本并大幅提升解题质量。案例研究显示,不同领域生成的工具可组合解决复合任务。综合来看,基于大模型的智能体能利用科学知识与编程能力自主构建可复用工具,指向一种以任务定义智能体能力的新范式。
原文摘要 · Abstract (English)
AI for science promises to accelerate the discovery process. The advent of large language models (LLMs) and agentic workflows enables the expediting of a growing range of scientific tasks. However, most of the current generation of agentic systems depend on static, hand-curated toolsets that hinder adaptation to new domains and evolving libraries. We present El Agente Forjador, a multi-agent framework in which universal coding agents autonomously forge, validate, and reuse computational tools through a four-stage workflow of tool analysis, tool generation, task execution, and iterative solution evaluation. Evaluated across 24 tasks spanning quantum chemistry and quantum dynamics on five coding agent setups, we compare three operating modes: zero-shot generation of tools per task, reuse of a curriculum-built toolset, and direct problem-solving with the coding agents as the baseline. We find that our tool generation and reuse framework consistently improves accuracy over the baseline. We also show that reusing a toolset built by a stronger coding agent can reduce API cost and substantially raises the solution quality for weaker coding agents. Case studies further demonstrate that tools forged for different domains can be combined to solve hybrid tasks. Taken together, these results show that LLM-based agents can use their scientific knowledge and coding capabilities to autonomously build reusable scientific tools, pointing toward a paradigm in which agent capabilities are defined by the tasks they are designed to solve rather than by explicitly engineered implementations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。