让代码像科学家一样思考,自动设计并修复复杂科学计算流程。
ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms
- 用专家知识构建概念框架,分步完成算法设计与纠错。
- 在标准测试中准确率超过或媲美人工编写的参考方案。
- 适合需要高可靠性的科学计算与机器学习研究者。
计算科学的进步依赖于复杂的数值工作流,必须精确体现物理规律,但将科学洞见转化为可靠代码仍是主要瓶颈。尽管大语言模型能生成孤立代码片段,却缺乏系统性推理能力来设计、验证和迭代优化完整科学管道。我们提出ATHENA,一种模拟科研过程的代理框架,将其建模为基于知识的上下文老虎机问题。其核心循环通过专家提供的概念支架,将概念策略与数值实现分离,实现对计算策略的系统诊断、重构与修复。在科学计算与科学机器学习任务中,ATHENA可自主推导并正确应用精确解析解,构建稳定数值求解器,诊断病态公式,并协调符号-数值混合工作流。定量结果显示,它在经典基准上达到甚至超过文献中专家编写的参考解的精度。通过将计算本身视为代理推理的对象,该框架实现了跨科学领域的异构算法自主调度。
原文摘要 · Abstract (English)
Progress in computational science depends on complex numerical workflows that must faithfully encode physical laws, yet translating conceptual insight into reliable code remains a major bottleneck. Although large language models can generate isolated code fragments, they lack the structured reasoning required to design, verify, and iteratively refine complete scientific pipelines. Here we introduce ATHENA, an agentic framework explicitly designed to emulate scientific research modeled as a knowledge-driven contextual bandit process. Its core loop separates conceptual policy from numerical realization through expert-derived conceptual scaffolding, enabling principled diagnosis, reformulation, and repair of computational strategies. Across scientific computing and scientific machine learning tasks, ATHENA autonomously derives and correctly applies exact analytical solutions, constructs stable numerical solvers, diagnoses ill-posed formulations, and orchestrates hybrid symbolic-numeric workflows. Quantitatively, ATHENA matches and frequently surpasses the accuracy of expert-authored reference solutions reported in the literature on canonical benchmarks. By reframing computation as an object of agentic reasoning, our framework enables autonomous orchestration of heterogeneous algorithms across scientific domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。