AI自主科研框架通过树状结构持续积累经验,实现长期迭代优化。
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

- 用树形结构关联假设、证据与洞见,支持跨时间累积学习。
- 在6项真实研究任务中均超越Codex和Claude Code,性能提升超2.5倍。
- 适合需要长期自主优化的AI科研系统开发者或研究者。
科学进步依赖于探索、实验与抽象的循环。研究人员测试候选方向,解读证据,并将经验应用于后续尝试。我们研究如何让AI代理在长周期内自主运行这一循环。提出Arbor框架,结合长期协调器、短期执行器与假设树精炼(HTR)机制——一种持久的树结构,连接假设、产物、证据与提炼出的洞见。协调器管理全局研究策略,执行器在隔离的工作树中实施并验证具体假设。结果返回后,Arbor更新树结构,传播可复用的经验,精炼搜索前沿,并接纳已验证的改进。该设计将自主科研从一系列局部尝试转化为可累积的过程,使策略、执行与证据得以跨时间延续。我们在自主优化(AO)设置下评估Arbor,即代理在无步骤级人类监督下,通过迭代实验优化初始研究产物。在模型训练、工具工程和数据合成等六项真实任务中,Arbor在所有任务上均取得最优保留结果,在相同接口与资源预算下,相对增益超过Codex与Claude Code平均值2.5倍以上。在MLE-Bench Lite上,Arbor使用GPT-5.5达到86.36%任意奖牌率,为对比中最佳表现。
原文摘要 · Abstract (English)
Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the evidence, and carry the resulting lessons into later attempts. We study how an AI agent can run this loop autonomously over long horizons. We introduce Arbor, a general framework for autonomous research that combines a long-lived coordinator, short-lived executors, and Hypothesis Tree Refinement (HTR), a persistent tree that links hypotheses, artifacts, evidence, and distilled insights across time. The coordinator manages global research strategy over the tree, while executors implement and test individual hypotheses in isolated worktrees. As results return, Arbor updates the tree, propagates reusable lessons, refines the search frontier, and admits verified improvements. This design turns autonomous research from a sequence of local attempts into a cumulative process in which strategy, execution, and evidence are carried across time. We evaluate Arbor under Autonomous Optimization (AO), an operational setting where an agent improves an initial research artifact through iterative experimentation without step-level human supervision. Across six real research tasks in model training, harness engineering, and data synthesis, Arbor achieves the best held-out result on all six tasks, attaining more than 2.5x the average relative held-out gain of Codex and Claude Code under the same task interface and resource budget. On MLE-Bench Lite, Arbor reaches 86.36% Any Medal with GPT-5.5, the strongest result in our comparison.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。