arXiv:2603.17216cs.AI2026-03

用自动生成的训练任务,让机器学习代理学会完整科研流程。

ML-AutoResearch: Training Machine Learning Research Agents with Automatically Generated Environments

  • 自动生成包含调试与迭代的全流程机器学习研究任务。
  • 在3个基准上提升9%的性能聚合指标,通过率显著提高。
  • 适合研究AI自动化科研、智能体训练的学者与工程师。

随着AI代理的发展,自动化科学发现正变得越来越可行。然而,训练代理自主完成机器学习研究中的工程密集型工作,需要大量过程级监督。现有静态基准忽略了调试和增量推理等关键中间步骤,而人工数据收集成本过高。为克服这一数据瓶颈,我们提出ML-AutoResearch(ML-AR),一个可扩展的自动合成端到端机器学习研究任务的流水线。每个任务定义完整的科研周期,包括问题设定、数据集选择、基线实现和迭代优化。为确保真实性和可执行性,任务基于真实数据集,并通过无需人工干预的自动化自调试流程进行优化。我们构建了一个大规模教师轨迹数据集,用于在这些合成任务上对学生代理进行监督微调。我们在3个不同的机器学习研究基准上评估了所生成的代理。跨2种模型族和3种模型规模的全面实验表明,在ML-AR轨迹上微调能带来持续且显著的能力提升。微调使性能聚合指标(AUP)最高提升9%,并大幅提高整体通过率,凸显出强大的跨领域泛化能力。

原文摘要 · Abstract (English)

With the advent of AI agents, automated scientific discovery is becoming an increasingly plausible goal. However, training agents to autonomously execute the engineering-heavy labor of machine learning (ML) research requires massive, process-level supervision. Existing static benchmarks omit critical intermediate steps such as debugging and incremental reasoning, and manual data collection is prohibitively expensive. To overcome this data bottleneck, we introduce ML-AutoResearch (ML-AR), a scalable pipeline for automatically generating synthetic, end-to-end ML research tasks. Each task defines a complete research cycle, including problem specification, dataset selection, baseline implementation, and iterative improvement. To ensure realism and executability, tasks are grounded in real-world datasets and refined via an automated self-debugging procedure without requiring human supervision. We construct a large-scale dataset of teacher trajectories on these synthetic tasks to train student agents via supervised fine-tuning. We evaluate the resulting agents across 3 diverse ML research benchmarks. Our comprehensive experiments across 2 distinct model families and 3 model sizes demonstrate that training on ML-AR trajectories yields consistent and significant capability gains. Fine-tuning improves the Aggregation Under the Performance (AUP) by up to 9\% and substantially boosts overall pass rates, highlighting robust out-of-domain generalization.

AI科研智能体训练自动化研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。