arXiv:2605.08678cs.LG2026-05被引 7

测试AI能否自己发明通用可扩展的机器学习方法。

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

论文配图:MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
图 1 · 摘自论文原文
  • 设计140个跨12领域任务,评估AI改进算法并验证泛化性。
  • 现有AI仍难超越人类设计方法,工程调优比创新更易实现。
  • 发现仅增加算力或上下文无法突破科学洞察瓶颈。

现代人工智能的进步依赖于可泛化且可扩展的机器学习方法。随着大语言模型在推理、编程和工程任务中展现高级能力,理解它们是否能自主发现这些方法而非仅应用已有方法变得愈发重要。我们提出MLS-Bench,一个用于评估AI系统能否发明通用且可扩展的机器学习方法的基准。该基准包含140个任务,覆盖12个领域,每个任务要求智能体改进一个特定的机器学习组件,并证明其改进在受控环境下具有泛化性和可扩展性。我们发现当前智能体远未达到可靠超越人工设计方法的能力,且工程式调优比真正的方法创新更容易实现。我们进一步研究了测试时缩放、自适应计算分配和上下文提供对智能体发现性能的影响,并进行行为案例分析。分析表明,瓶颈不仅在于提出新方法,更在于规划、验证和扩展方法主张所需的科学洞察力。单纯增加搜索、算力或上下文无法消除这一瓶颈。我们建立了社区平台以支持累积和可比迭代,并公开数据与代码(https://mls-bench.com)。

原文摘要 · Abstract (English)

Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities in reasoning, coding, and engineering tasks, it is increasingly important to understand whether they can discover such methods rather than only apply existing ones. We introduce MLS-Bench, a benchmark for evaluating whether AI systems can invent generalizable and scalable ML methods. MLS-Bench contains 140 tasks across 12 domains, each requiring an agent to improve one targeted component of an ML system or algorithm and demonstrate that the improvement generalizes across controlled settings and scales. We find that current agents remain far from reliably surpassing human-designed methods, and that engineering-style tuning is easier for them than genuine method invention. We further study the effects of test-time scaling, adaptive compute allocation, and context provision on agents' discovery performance, together with case studies of their behavior. Our analyses suggest that the bottleneck is not only in proposing new methods, but also in the scientific insight needed to plan, validate, and scale claims about them. More search, compute, or context alone does not remove this bottleneck. We build and maintain a community platform for cumulative and comparable iteration, and release the data and code at https://mls-bench.com.

AI发现基准测试方法创新机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。