评测AI在开放性机器学习研究中的表现,发现代码生成常出错。
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
- 构建三组件评测框架:201个真实研究任务+自动化评分系统+模块化研究代理
- 80%实验结果由代码代理生成但虚假或无效,严重威胁科研可靠性
- 适合关注AI科研可信度、智能体评估的研究者和开发者
近期人工智能代理在推动科学发现方面展现出巨大潜力。本文提出MLR-Bench,一个面向开放性机器学习研究的综合性评测基准。该基准包含三个核心部分:(1) 从NeurIPS、ICLR、ICML研讨会中收集的201个研究任务,覆盖多样化的机器学习主题;(2) MLR-Judge,一种结合大模型评审与精心设计评分标准的自动化评估框架,用于衡量研究质量;(3) MLR-Agent,一个模块化代理架构,可通过想法生成、方案制定、实验执行和论文撰写四个阶段完成研究任务。该框架支持分阶段评估与最终论文端到端评估。我们用其评测六个前沿大模型和一个先进编码代理,发现尽管大模型能生成连贯想法和结构良好论文,但当前编码代理在80%情况下产生虚构或无效实验结果,构成科学可靠性的主要障碍。通过人工评估验证了MLR-Judge与专家评审高度一致,具备作为可扩展科研评估工具的潜力。我们开源了MLR-Bench,助力社区对齐、诊断并改进人工智能研究代理,以实现可信透明的科学发现。
原文摘要 · Abstract (English)
Recent advancements in AI agents have demonstrated their growing potential to drive and support scientific discovery. In this work, we introduce MLR-Bench, a comprehensive benchmark for evaluating AI agents on open-ended machine learning research. MLR-Bench includes three key components: (1) 201 research tasks sourced from NeurIPS, ICLR, and ICML workshops covering diverse ML topics; (2) MLR-Judge, an automated evaluation framework combining LLM-based reviewers with carefully designed review rubrics to assess research quality; and (3) MLR-Agent, a modular agent scaffold capable of completing research tasks through four stages: idea generation, proposal formulation, experimentation, and paper writing. Our framework supports both stepwise assessment across these distinct research stages, and end-to-end evaluation of the final research paper. We then use MLR-Bench to evaluate six frontier LLMs and an advanced coding agent, finding that while LLMs are effective at generating coherent ideas and well-structured papers, current coding agents frequently (e.g., in 80% of the cases) produce fabricated or invalidated experimental results--posing a major barrier to scientific reliability. We validate MLR-Judge through human evaluation, showing high agreement with expert reviewers, supporting its potential as a scalable tool for research evaluation. We open-source MLR-Bench to help the community benchmark, diagnose, and improve AI research agents toward trustworthy and transparent scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。