测试大模型能否解决真实机器学习难题,发现当前能力仍有巨大差距。
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
- 设计7个真实科研挑战任务,评估模型提出并实现新方法的能力。
- 最佳模型仅弥补人类顶尖水平9.3%的分数差距。
- 揭示模型自评创新与实际性能间的严重不符,适合研究AI科研能力者参考。
我们提出了MLRC-Bench,一个用于量化语言代理在解决复杂机器学习研究竞赛中表现的基准,重点关注需要新颖方法论的开放性研究问题。与以往工作(如AI Scientist)不同,后者通过大模型作为裁判评估端到端代理流程,MLRC-Bench则衡量提出和实现新研究方法的关键步骤,并采用严格协议与客观指标进行评估。我们精心筛选的7个竞赛任务揭示了语言代理面临的显著挑战:即使表现最好的测试代理(gemini-exp-1206,在MLAB设置下)也仅能缩小基线与顶尖人类参与者之间9.3%的分数差距。此外,我们的分析表明,大模型自我评判的创新性与其在前沿机器学习问题上的实际表现存在明显错位。MLRC-Bench是一个动态基准,旨在随新机器学习竞赛持续演进,并推动对AI研究能力的严谨、客观评估。排行榜与代码已公开于:https://huggingface.co/spaces/launch/MLRC_Bench
原文摘要 · Abstract (English)
We introduce MLRC-Bench, a benchmark designed to quantify how effectively language agents can tackle challenging Machine Learning (ML) Research Competitions, with a focus on open research problems that demand novel methodologies. Unlike prior work, e.g., AI Scientist, which evaluates the end-to-end agentic pipeline by using LLM-as-a-judge, MLRC-Bench measures the key steps of proposing and implementing novel research methods and evaluates them with rigorous protocol and objective metrics. Our curated suite of 7 competition tasks reveals significant challenges for LLM agents. Even the best-performing tested agent (gemini-exp-1206 under MLAB) closes only 9.3% of the gap between baseline and top human participant scores. Furthermore, our analysis reveals a misalignment between the LLM-judged innovation and actual performance on cutting-edge ML research problems. MLRC-Bench is a dynamic benchmark, designed to grow with new ML competitions and encourage rigorous, objective evaluations of AI research capabilities. Our leaderboard and code are available at: https://huggingface.co/spaces/launch/MLRC_Bench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。