arXiv:2510.11694cs.AI2025-10被引 2

一个智能体搞定机器学习全流程,效率超越多智能体系统。

Operand Quant: A Single-Agent Architecture for Autonomous Machine Learning Engineering

  • 用单个智能体在集成开发环境里完成从实验到部署的全部流程。
  • 在MLE-Benchmark上达成0.3956的综合奖牌率,创历史新高。
  • 适合追求高效自动化机器学习工程的研究者和开发者。

我们提出Operand Quant,一种基于集成开发环境的单智能体自治机器学习工程架构。该架构将机器学习全生命周期——探索、建模、实验与部署——整合于一个具备上下文感知能力的单一智能体中,摒弃传统多智能体协同框架。在MLE-Benchmark(2025)测试中,Operand Quant在75个问题上取得0.3956±0.0565的综合奖牌率,为当前所有参评系统中的最高纪录。结果表明,在相同约束条件下,一个线性、非阻塞的自主智能体可在受控IDE环境中超越多智能体协作系统的表现。

原文摘要 · Abstract (English)

We present Operand Quant, a single-agent, IDE-based architecture for autonomous machine learning engineering (MLE). Operand Quant departs from conventional multi-agent orchestration frameworks by consolidating all MLE lifecycle stages -- exploration, modeling, experimentation, and deployment -- within a single, context-aware agent. On the MLE-Benchmark (2025), Operand Quant achieved a new state-of-the-art (SOTA) result, with an overall medal rate of 0.3956 +/- 0.0565 across 75 problems -- the highest recorded performance among all evaluated systems to date. The architecture demonstrates that a linear, non-blocking agent, operating autonomously within a controlled IDE environment, can outperform multi-agent and orchestrated systems under identical constraints.

机器学习自动化智能体IDE集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。