用多智能体系统实现长期机器学习研究的自动化工程。
Toward Autonomous Long-Horizon Engineering for ML Research
- 构建轻量级协作框架,通过文件即总线机制保存项目状态。
- 在PaperBench和MLE-Bench上分别提升9.92至16.67分,超越基线。
- 适合需要长期迭代、可追溯实验过程的研究者使用。
智能体系统正逐步自动化人工智能研究的局部环节,但将模糊的研究目标转化为可运行、经实验验证的机器学习系统仍是核心瓶颈。本文将此场景定义为‘长时序机器学习研究工程’:通过反复实现、实验与优化,将研究需求转化为可运行的机器学习系统。核心挑战在于在异构阶段中维持累积性进展,同时应对延迟且混杂的反馈。为此提出AiScientist,一种围绕轻量控制、厚态状态构建的多智能体系统:由轻量级层级研究团队协作,通过‘文件即总线’工作区保留跨角色与调用的关键决策产物。在PaperBench上,AiScientist使用Gemini-3-Flash和GLM-5分别比最强基线提升9.92和11.15分;在MLE-Bench Lite上,两种骨干模型均达81.82% Any Medal%,分别超越基线4.55和16.67分,并领先于Codex/GPT-5.5 xhigh frontier参考水平13.64分。消融实验与过程分析表明,持久项目状态对后期优化至关重要:移除文件即总线使PaperBench得分下降6.41分,MLE-Bench Lite Any Medal%下降31.82分。结果表明,长时序AI研究不仅是局部推理能力问题,更是维持可追溯、可持续进展的系统工程问题。
原文摘要 · Abstract (English)
Agentic systems increasingly automate pieces of AI research. Yet turning underspecified research objectives into runnable, experimentally validated ML systems remains a central bottleneck. We study this operational setting as \emph{long-horizon ML research engineering}: converting a research specification into a runnable ML system through repeated implementation, experimentation, and refinement. The central challenge is to sustain cumulative project progress across heterogeneous stages under delayed, confounded feedback. We introduce AiScientist, a multi-agent system built around thin control over thick state: a lightweight hierarchical research team coordinates through a File-as-Bus workspace that preserves decision-relevant artifacts across roles and invocations. On PaperBench, AiScientist improves over the strongest matched baselines by 9.92 and 11.15 points with Gemini-3-Flash and GLM-5, respectively. On MLE-Bench Lite, it reaches 81.82 Any Medal\% under both backbones, improving over the strongest matched baselines by 4.55 and 16.67 points, and exceeding a Codex/GPT-5.5 xhigh frontier harness reference by 13.64 Any Medal points. Ablations and process analyses show that durable project state is central to later-round refinement: removing File-as-Bus lowers PaperBench score by 6.41 points and MLE-Bench Lite Any Medal\% by 31.82 points. These results suggest that long-horizon AI research is not only a problem of stronger local reasoning, but a systems problem of maintaining cumulative, inspectable project progress.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。