arXiv:2502.00964cs.SEcs.AI2025-02被引 17

测试AI代理在真实机器学习开发流程中的综合能力。

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows

  • 设计涵盖数据处理到模型部署的30个全流程任务
  • 三款代理在调试与工具集成上表现差异明显
  • 适合关注AI工程化落地的研究者与开发者

本文提出ML-Dev-Bench,一个针对实际机器学习开发流程中智能体能力的基准测试。不同于聚焦单一编码或竞赛任务的现有基准,该基准评估代理在数据处理、模型训练、改进已有模型、调试及与主流机器学习工具集成等方面的综合表现。我们在30个多样化任务上评测了ReAct、Openhands和AIDE三个代理,揭示其在应对真实开发挑战时的优势与局限。相关代码已开源,供社区使用。

原文摘要 · Abstract (English)

In this report, we present ML-Dev-Bench, a benchmark aimed at testing agentic capabilities on applied Machine Learning development tasks. While existing benchmarks focus on isolated coding tasks or Kaggle-style competitions, ML-Dev-Bench tests agents' ability to handle the full complexity of ML development workflows. The benchmark assesses performance across critical aspects including dataset handling, model training, improving existing models, debugging, and API integration with popular ML tools. We evaluate three agents - ReAct, Openhands, and AIDE - on a diverse set of 30 tasks, providing insights into their strengths and limitations in handling practical ML development challenges. We open source the benchmark for the benefit of the community at \href{https://github.com/ml-dev-bench/ml-dev-bench}{https://github.com/ml-dev-bench/ml-dev-bench}.

AI代理机器学习基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。