构建可自主完成数据科学任务的智能体框架,提升机器学习工程效率。
R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science
- 将机器学习工程拆解为两阶段六组件,实现流程化与可测试化
- 在MLE-Bench上达成35.1%任意奖牌率,领先现有方法
- 适合研究者与工程师快速构建高效自动化机器学习系统
人工智能与机器学习的进展虽已改变数据科学,但复杂度与专业门槛仍制约发展。尽管众包平台缓解部分挑战,高阶机器学习工程任务仍耗时且迭代频繁。本文提出R&D-Agent,一个全面、解耦且可扩展的框架,形式化机器学习工程流程。该框架将流程划分为两个阶段、六个组件,使机器学习工程智能体设计从零散经验转变为有原则、可验证的过程。尽管现有智能体在特定组件上表现良好,但多数可归结为本框架基线的局部优化。受人类专家启发,我们在该框架内设计出高效智能体,达到业界顶尖性能。在MLE-Bench评测中,基于R&D-Agent的智能体位居榜首,任意奖牌率达35.1%,证明该框架能显著加速创新并提升数据科学应用的准确性。项目已开源:https://github.com/microsoft/RD-Agent。
原文摘要 · Abstract (English)
Recent advances in AI and ML have transformed data science, yet increasing complexity and expertise requirements continue to hinder progress. Although crowd-sourcing platforms alleviate some challenges, high-level machine learning engineering (MLE) tasks remain labor-intensive and iterative. We introduce R&D-Agent, a comprehensive, decoupled, and extensible framework that formalizes the MLE process. R&D-Agent defines the MLE workflow into two phases and six components, turning agent design for MLE from ad-hoc craftsmanship into a principled, testable process. Although several existing agents report promising gains on their chosen components, they can mostly be summarized as a partial optimization from our framework's simple baseline. Inspired by human experts, we designed efficient and effective agents within this framework that achieve state-of-the-art performance. Evaluated on MLE-Bench, the agent built on R&D-Agent ranks as the top-performing machine learning engineering agent, achieving 35.1% any medal rate, demonstrating the ability of the framework to speed up innovation and improve accuracy across a wide range of data science applications. We have open-sourced R&D-Agent on GitHub: https://github.com/microsoft/RD-Agent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。