构建首个面向机器学习研究级任务的基准,评估AI代理真实科研能力。
ML Research Benchmark
- 基于顶会竞赛题设计7个真实研究任务,覆盖模型训练、微调与压缩等环节。
- Claude-3.5 Sonnet在规划和建模上表现最优,但无法完成复杂研究迭代。
- 揭示当前大模型代理在高级科研中仍存短板,适合研究评估工具开发者参考。
人工智能代理在多个领域执行复杂任务的能力不断提升。随着其发展,亟需精准衡量和评估其能力,尤其在加速人工智能研究与开发方面。现有基准主要关注通用机器学习任务,缺乏对AI代理应对研究级问题及竞赛级挑战的全面评估方法。本文提出机器学习研究基准(MLRB),包含7个源自近期机器学习顶会赛道的竞赛级任务,涵盖AI研究人员典型工作,如模型训练效率、有限数据预训练、领域特定微调和模型压缩。本研究引入由前沿模型(包括Claude-3和GPT-4o)驱动的代理框架进行评估。结果表明,Claude-3.5 Sonnet代理在基准上表现最佳,尤其在规划与模型开发方面优势明显;然而,所有测试代理均难以完成非平凡的研究迭代。任务间表现差异显著,凸显了人工智能研发的复杂性以及构建通用代理框架的挑战。尽管当前代理能有效理解复杂指令并生成基线结果,但仍远未达到先进人工智能研究所需的能力。MLRB为评估和比较代理在模拟真实世界人工智能研究挑战中的表现提供了宝贵框架。
原文摘要 · Abstract (English)
Artificial intelligence agents are increasingly capable of performing complex tasks across various domains. As these agents advance, there is a growing need to accurately measure and benchmark their capabilities, particularly in accelerating AI research and development. Current benchmarks focus on general machine learning tasks, but lack comprehensive evaluation methods for assessing AI agents' abilities in tackling research-level problems and competition-level challenges in the field of AI. We present the ML Research Benchmark (MLRB), comprising 7 competition-level tasks derived from recent machine learning conference tracks. These tasks span activities typically undertaken by AI researchers, including model training efficiency, pretraining on limited data, domain specific fine-tuning, and model compression. This paper introduces a novel benchmark and evaluates it using agent scaffolds powered by frontier models, including Claude-3 and GPT-4o. The results indicate that the Claude-3.5 Sonnet agent performs best across our benchmark, excelling in planning and developing machine learning models. However, both tested agents struggled to perform non-trivial research iterations. We observed significant performance variations across tasks, highlighting the complexity of AI development and the challenges in creating versatile agent scaffolds. While current AI agents can successfully navigate complex instructions and produce baseline results, they fall short of the capabilities required for advanced AI research. The ML Research Benchmark provides a valuable framework for assessing and comparing AI agents on tasks mirroring real-world AI research challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。