arXiv:2502.14499cs.CLcs.AI2025-02被引 100

首个面向大模型科研能力的评估框架,支持强化学习训练与多领域任务测试。

MLGym: A New Framework and Benchmark for Advancing AI Research Agents

  • 构建首个机器学习任务的Gym环境,支持强化学习训练智能体。
  • 涵盖13个跨领域开放任务,需完成从实验设计到迭代优化全流程。
  • 当前顶尖模型仅能调优参数,无法生成新算法或突破性创新,适合研究者参考。

我们推出Meta MLGym与MLGym-Bench,首个面向机器学习任务的大语言模型智能体评估框架与基准。该框架首次实现机器学习任务的Gym环境化,支持强化学习算法训练智能体。MLGym-Bench包含13个来自计算机视觉、自然语言处理、强化学习和博弈论等领域的多样化、开放性人工智能研究任务。解决这些任务需具备提出新假设、数据处理、模型实现、实验训练、结果分析及迭代优化等真实科研能力。我们在多个前沿大模型(如Claude-3.5-Sonnet、Llama-3.1 405B、GPT-4o、o1-preview、Gemini-1.5 Pro)上进行评估,发现当前模型通常通过优化超参数提升基线表现,但未能生成新理论、新算法或架构,也未带来显著性能突破。框架支持任务扩展、模型集成、大规模合成数据生成及新训练算法开发。我们已开源该框架与基准,以推动大模型科研能力研究。

原文摘要 · Abstract (English)

We introduce Meta MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks. This is the first Gym environment for machine learning (ML) tasks, enabling research on reinforcement learning (RL) algorithms for training such agents. MLGym-bench consists of 13 diverse and open-ended AI research tasks from diverse domains such as computer vision, natural language processing, reinforcement learning, and game theory. Solving these tasks requires real-world AI research skills such as generating new ideas and hypotheses, creating and processing data, implementing ML methods, training models, running experiments, analyzing the results, and iterating through this process to improve on a given task. We evaluate a number of frontier large language models (LLMs) on our benchmarks such as Claude-3.5-Sonnet, Llama-3.1 405B, GPT-4o, o1-preview, and Gemini-1.5 Pro. Our MLGym framework makes it easy to add new tasks, integrate and evaluate models or agents, generate synthetic data at scale, as well as develop new learning algorithms for training agents on AI research tasks. We find that current frontier models can improve on the given baselines, usually by finding better hyperparameters, but do not generate novel hypotheses, algorithms, architectures, or substantial improvements. We open-source our framework and benchmark to facilitate future research in advancing the AI research capabilities of LLM agents.

智能体强化学习科研自动化评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。