arXiv:2601.17596cs.CL2026-01Conference of the …被引 5

让AI分工协作,先想策略再执行,提升机器学习优化效率。

Learning to Ideate for Machine Learning Engineering Agents

  • 分离创意与执行,由专门代理提出优化策略
  • 仅用1000样本训练,效果比未训练版本高11.5%
  • 适合研究AI自主科研、自动化机器学习的团队

现有机器学习工程(MLE)代理在迭代优化算法时表现不佳。为此,我们提出MLE-Ideator,一种双代理框架,将创意生成与实现分离。实现代理可向专门的创意代理请求策略支持。实验表明,该方法有效:一是在无训练设置下,显著优于仅实现型代理基线;二是创意代理可通过强化学习训练,生成更优策略。仅需10个任务的1000个训练样本,基于Qwen3-8B的训练后创意代理相比未训练版本实现11.5%相对提升,并超越Claude Sonnet 3.5。结果展示了训练战略性AI系统用于科学发现的潜力。

原文摘要 · Abstract (English)

Existing machine learning engineering (MLE) agents struggle to iteratively optimize their implemented algorithms for effectiveness. To address this, we introduce MLE-Ideator, a dual-agent framework that separates ideation from implementation. In our system, an implementation agent can request strategic help from a dedicated Ideator. We show this approach is effective in two ways. First, in a training-free setup, our framework significantly outperforms implementation-only agent baselines on MLE-Bench. Second, we demonstrate that the Ideator can be trained with reinforcement learning (RL) to generate more effective ideas. With only 1K training samples from 10 MLE tasks, our RL-trained Qwen3-8B Ideator achieves an 11.5% relative improvement compared to its untrained counterpart and surpasses Claude Sonnet 3.5. These results highlights a promising path toward training strategic AI systems for scientific discovery.

机器学习AI代理强化学习自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。