arXiv:2608.14354cs.AI2026-08

让AI长期自主研究,突破传统智能体的探索瓶颈。

ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond

论文配图:ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond
图 1 · 摘自论文原文
  • 将复杂研究拆解为可执行、可恢复的阶段性任务
  • 24小时内达成MLBench 70.22%的最高综合得分
  • 适合需要持续攻关的科研自动化场景

让大模型智能体在长时间跨度内保持高效、稳定且目标一致的研究能力,是实现自主机器学习与科学发现的核心挑战。现有自研智能体虽取得进展,但仍缺乏连续性管理、死胡同恢复机制及价值驱动的算力分配策略,导致搜索效率低下、资源浪费且成功率降低。为此,我们提出ScienceFlow——一个端到端的自研智能体框架,将长周期研究分解为基于可执行工作区的研究段。通过可恢复的可执行状态表示研究进展,支持高效探索、修正与执行。研究段间的转换由可执行状态重锚转移(ESTRA)机制控制,动态选择当前状态或归档状态作为新锚点,并决定是否继续或调整研究路径。证据感知执行控制器根据资源可用性、剩余预算和验证进度分配算力。我们在机器学习、科学建模与数学优化等任务上评估了ScienceFlow。在多个长周期基准测试中表现优异,尤其在24小时预算下于MLE-bench全量评测中达到70.22%的任意奖牌率,领先先前结果4.92个百分点。结果表明,高效的态管理、自适应探索与目标对齐的执行,是拓展自主研究至长周期任务的关键。

原文摘要 · Abstract (English)

Enabling LLM agents to sustain productive, stable, and goal-aligned research over extended horizons is a central challenge for autonomous machine learning and scientific discovery, as progress hinges on continuously managing evolving state, exploration decisions, and computational resources. Pioneering autoresearch agents, despite great success, still lack mechanisms for continuity, recovery from dead ends, and value-driven compute allocation, which inherently undermines overall search efficiency, wastes computational resources, and lowers the chance of ultimate success. To bridge this gap, we introduce ScienceFlow, an end-to-end autoresearch agent framework that organizes long-horizon research work into research segments grounded in executable workspaces. It represents research progress as recoverable executable states, enabling efficient exploration, revision, and execution. Transitions between research segments are governed by Executable-State Transition through Re-Anchoring (ESTRA), which selects either the live state or an archived state as the next anchor and determines whether to continue or redirect the research trajectory. An evidence-aware execution controller allocates resources to physical jobs based on resource availability, remaining budget, and validated progress. We evaluate ScienceFlow on tasks spanning machine learning, scientific modeling, and mathematical optimization. Results on diverse long-horizon benchmarks demonstrate its ability to sustain effective research processes, highlighted by a SOTA 70.22 percent Any-Medal score on the full MLE-bench within a 24-hour budget, outperforming prior reported results by 4.92 percentage points. The efficacy of ScienceFlow further demonstrates that efficient state management, adaptive exploration, and objective-aligned execution are critical for scaling autonomous research beyond short-horizon interactions.

自主科研智能体长程任务研究自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。