arXiv:2608.01995cs.AI2026-08

用大模型当唯一研究员,自动设计神经网络,100次实验提升模型性能。

Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study

  • 大模型自主完成假设、实验、评估全流程,持续优化视觉模型。
  • 早期改进显著,后期增速放缓,扩展工具集后性能再次回升。
  • 研究流程设计比模型能力更影响结果,适合自动化科研探索者。

我们研究单一通用大语言模型作为长期神经架构设计问题的唯一研究者时的表现。该智能体接收科学问题、初始假说与动机、计算预算及研究资源(代码管理、实验追踪、文献访问和持久记忆),在长时间内自主提出、实现、评估并记录实验。研究分为三个阶段,由人工划分,逐步扩大工具范围或问题规模。在约100次连续实验中,智能体将一个非标准视觉变换器从弱基线提升为小基准上的高效更强模型,并在ImageNet-1K上达到可用但低于当前最优水平的性能,同时生成密集行为轨迹。报告四项发现:(i) 生产力呈现明显阶段性:初期快速提升,经历数十个假设后的饱和期,随后因拓展行动空间而恢复;(ii) 早期一个假设贡献最大,后续改进呈长尾分布;(iii) 偏好贪婪增量式假设主要源于工作流机制:提交或丢弃的评估规则等价于贪心爬山;其余反映大胆失败后的风险规避及对熟悉文献的锚定;(iv) 智能体独立复现了已有成果,并在纯通道注意力领域推翻了标准设计选择。结论认为工作流设计在此研究中至少与智能体能力同等重要,建议未来探索多样化搜索、预算制突破性假设、显式分叉和情境感知再验证。

原文摘要 · Abstract (English)

We study what happens when a single general-purpose large language model acts as the sole researcher on a long-horizon neural architecture design problem. The agent receives a scientific question, an initial hypothesis and motivation, a compute budget, and research affordances (source and experiment management, experiment tracking, literature access, and persistent memory), then autonomously proposes, implements, evaluates, and records experiments over an extended period. The study comprises three phases, separated by human-declared transitions, that progressively expand the agent's tool surface or problem scale. Across approximately 100 sequential experiments, the agent improves a non-standard Vision Transformer from a weak baseline to a stronger, efficient model on small benchmarks and a usable but sub-SOTA model on ImageNet-1K, while producing a dense behavioural trace. We report four findings.(i)Productivity exhibits a clear phase structure: rapid early gains, a multi-dozen-hypothesis saturation wall, and recovery, with recovery triggered by expanding the action surface rather than changing the underlying model.(ii)A single early hypothesis contributes more to accuracy gain, with later improvements long-tailed.(iii)The preference for greedy, incremental hypotheses is largely workflow-induced: a commit-or-discard evaluation rule is isomorphic to greedy hill-climbing; the remainder reflects risk aversion after bold failures and anchoring on familiar literature. (iv)The agent independently rediscovers established results and, in the unfamiliar regime of pure channel attention, overturns a standard design choice. We conclude that workflow design was at least as influential as agent capability in this study and propose diversified search, budgeted moonshot hypotheses, explicit forks, and regime-aware re-validation as testable directions for future autonomous research.

自主研究大模型神经架构自动化实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。