arXiv:2507.02554cs.AIcs.LG2025-07NeurIPS被引 58

AI研究代理在真实机器学习竞赛中实现更高胜率,关键在于搜索策略与操作集的协同优化。

AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench

  • 将AI研究代理视为搜索策略,在候选解空间中迭代修改解决方案。
  • 最优组合使MLE-bench lite上获奖成功率从39.6%提升至47.7%。
  • 适用于自动化机器学习、智能科研助手研发者参考。

AI研究代理在自动化机器学习模型设计、实现和训练方面展现出巨大潜力,有望加速科学进展。本文聚焦于提升代理在MLE-bench这一挑战性基准上的表现,该基准要求代理参与类似Kaggle的真实世界机器学习竞赛。我们把AI研究代理形式化为在候选解空间中导航的搜索策略,并通过迭代使用操作符进行修改。通过系统设计并调整不同操作符集合与搜索策略(贪心、蒙特卡洛树搜索、进化算法),发现其相互作用对性能至关重要。最优的搜索策略与操作符组合在MLE-bench lite上取得当前最优结果,将获得Kaggle奖牌的成功率从39.6%提升至47.7%。研究强调了在推进自动化机器学习时,需协同考虑搜索策略、操作符设计与评估方法。

原文摘要 · Abstract (English)

AI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus on methods for improving agents' performance on MLE-bench, a challenging benchmark where agents compete in Kaggle competitions to solve real-world machine learning problems. We formalize AI research agents as search policies that navigate a space of candidate solutions, iteratively modifying them using operators. By designing and systematically varying different operator sets and search policies (Greedy, MCTS, Evolutionary), we show that their interplay is critical for achieving high performance. Our best pairing of search strategy and operator set achieves a state-of-the-art result on MLE-bench lite, increasing the success rate of achieving a Kaggle medal from 39.6% to 47.7%. Our investigation underscores the importance of jointly considering the search strategy, operator design, and evaluation methodology in advancing automated machine learning.

AI研究代理自动化机器学习强化学习智能科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。