arXiv:2601.10402cs.AI2026-01被引 23

让AI在复杂科研中持续数天自主实验,突破传统模型局限。

Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering

  • 用分层认知缓存机制,将短期执行记录提炼为长期策略知识。
  • 在24小时任务中达成56.44%的奖牌率,超越现有水平。
  • 适合追求长周期自主科研的智能体开发者或研究者参考。

当前人工智能向自主科学迈进受限于超长时程自主能力,即在持续数日或数周的实验周期中保持战略连贯性与迭代修正能力。尽管大语言模型在短周期推理中表现优异,但在真实科研高维、延迟反馈环境中,易被执行细节淹没,难以将稀疏反馈整合为连贯的长期指导。本文提出ML-Master 2.0,一个能掌握超长时程机器学习工程(MLE)的自主智能体,该任务是科学发现的典型缩影。通过将上下文管理重构为认知积累过程,引入受计算机系统启发的分层认知缓存(HCC)架构,实现经验随时间的结构化分化。动态地将瞬时执行轨迹提炼为稳定知识与跨任务智慧,使代理能够解耦即时执行与长期实验策略,有效突破静态上下文窗口的扩展瓶颈。在OpenAI的MLE-Bench测试中,24小时预算下达到56.44%的奖牌率,创历史新高。结果表明,超长时程自主为超越人类先验复杂度的AI自主探索提供了可扩展范式。

原文摘要 · Abstract (English)

The advancement of artificial intelligence toward agentic science is currently bottlenecked by the challenge of ultra-long-horizon autonomy, the ability to sustain strategic coherence and iterative correction over experimental cycles spanning days or weeks. While Large Language Models (LLMs) have demonstrated prowess in short-horizon reasoning, they are easily overwhelmed by execution details in the high-dimensional, delayed-feedback environments of real-world research, failing to consolidate sparse feedback into coherent long-term guidance. Here, we present ML-Master 2.0, an autonomous agent that masters ultra-long-horizon machine learning engineering (MLE) which is a representative microcosm of scientific discovery. By reframing context management as a process of cognitive accumulation, our approach introduces Hierarchical Cognitive Caching (HCC), a multi-tiered architecture inspired by computer systems that enables the structural differentiation of experience over time. By dynamically distilling transient execution traces into stable knowledge and cross-task wisdom, HCC allows agents to decouple immediate execution from long-term experimental strategy, effectively overcoming the scaling limits of static context windows. In evaluations on OpenAI's MLE-Bench under 24-hour budgets, ML-Master 2.0 achieves a state-of-the-art medal rate of 56.44%. Our findings demonstrate that ultra-long-horizon autonomy provides a scalable blueprint for AI capable of autonomous exploration beyond human-precedent complexities.

自主智能体长时程推理机器学习工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。