arXiv:2604.05550cs.CLcs.CE2026-04被引 16

AutoSOTA自动发现并优化新模型,平均每篇论文5小时达成更优结果。

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery

论文配图:AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery
图 1 · 摘自论文原文
  • 多智能体协作,从论文到代码全流程自动化
  • 在8个顶会论文上发现105个新SOTA模型
  • 可识别架构创新与流程改进,适合研究者提速实验

人工智能研究日益依赖反复复现、调试和迭代优化以达到最先进性能,亟需加速完整模型优化流程的系统。本文提出AutoSOTA,一个端到端自动化研究系统,将顶级期刊论文中的最新SOTA模型转化为可复现且性能更优的新SOTA模型。该问题被分解为三个紧密耦合阶段:资源准备与目标设定、实验评估、反思与创意生成。AutoSOTA采用包含八个专业智能体的多智能体架构,协同完成论文到代码与依赖项的对齐、执行环境初始化与修复、长周期实验追踪、优化思路生成与调度,以及有效性监督以避免虚假提升。我们在来自八场顶级AI会议、满足代码可用性与执行成本限制的论文上评估AutoSOTA。结果显示,该系统在自动化复现与后续优化中均表现优异,成功发现105个超越原始方法的新SOTA模型,平均耗时约每篇5小时。涵盖大语言模型、自然语言处理、计算机视觉、时间序列与优化等领域的案例研究进一步表明,系统不仅限于超参调优,还能识别架构创新、算法重构与工作流优化。结果表明,端到端研究自动化不仅能作为性能优化器,还可作为新型科研基础设施,减轻重复实验负担,释放人类注意力于更高层次的科学创造力。

原文摘要 · Abstract (English)

Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creating a growing need for systems that can accelerate the full pipeline of empirical model optimization. In this work, we introduce AutoSOTA, an end-to-end automated research system that advances the latest SOTA models published in top-tier AI papers to reproducible and empirically improved new SOTA models. We formulate this problem through three tightly coupled stages: resource preparation and goal setting; experiment evaluation; and reflection and ideation. To tackle this problem, AutoSOTA adopts a multi-agent architecture with eight specialized agents that collaboratively ground papers to code and dependencies, initialize and repair execution environments, track long-horizon experiments, generate and schedule optimization ideas, and supervise validity to avoid spurious gains. We evaluate AutoSOTA on recent research papers collected from eight top-tier AI conferences under filters for code availability and execution cost. Across these papers, AutoSOTA achieves strong end-to-end performance in both automated replication and subsequent optimization. Specifically, it successfully discovers 105 new SOTA models that surpass the original reported methods, averaging approximately five hours per paper. Case studies spanning LLM, NLP, computer vision, time series, and optimization further show that the system can move beyond routine hyperparameter tuning to identify architectural innovation, algorithmic redesigns, and workflow-level improvements. These results suggest that end-to-end research automation can serve not only as a performance optimizer, but also as a new form of research infrastructure that reduces repetitive experimental burden and helps redirect human attention toward higher-level scientific creativity.

自动化研究SOTA发现多智能体模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。