arXiv:2507.18074cs.AI2025-07被引 21

AI自主发现106种新型注意力架构,突破人类设计瓶颈

AlphaGo Moment for Model Architecture Discovery

  • AI自主提出新架构构想并验证,实现从优化到创新的范式跃迁
  • 2万小时GPU算力下完成1773次实验,发现106个领先于人类的设计
  • 揭示可计算扩展的科学发现规律,为自加速AI研究提供蓝图

尽管人工智能能力呈指数级提升,但其研发速度仍受限于人类认知能力,形成严重瓶颈。我们提出ASI-Arch,首个在神经网络架构发现领域实现人工智能对人工智能(ASI4AI)的示范系统——一个完全自主的系统,能突破这一根本限制,实现架构自主创新。不同于传统神经架构搜索(NAS)仅在人类定义的空间内探索,该系统实现了从自动化优化到自动化创新的范式转变。ASI-Arch可端到端开展架构发现领域的科学研究:自主提出新颖架构概念、生成可执行代码、通过训练与实验验证性能,并利用过往经验迭代改进。系统累计投入20,000 GPU小时,完成1,773次自主实验,最终发现106种具备最先进水平(SOTA)的线性注意力架构。这些由AI发现的架构展现出超越人类设计基线的系统性优势,揭示了此前未知的设计路径,如同AlphaGo的第37手般带来颠覆性洞见。关键的是,我们首次建立科学发现本身的可扩展规律,证明架构突破可通过计算资源扩展实现,使研究进程从人力制约转向计算可扩展。我们对涌现设计模式与自主研究能力进行深入分析,为自加速人工智能系统提供可复现的蓝图。

原文摘要 · Abstract (English)

While AI systems demonstrate exponentially improving capabilities, the pace of AI research itself remains linearly bounded by human cognitive capacity, creating an increasingly severe development bottleneck. We present ASI-Arch, the first demonstration of Artificial Superintelligence for AI research (ASI4AI) in the critical domain of neural architecture discovery--a fully autonomous system that shatters this fundamental constraint by enabling AI to conduct its own architectural innovation. Moving beyond traditional Neural Architecture Search (NAS), which is fundamentally limited to exploring human-defined spaces, we introduce a paradigm shift from automated optimization to automated innovation. ASI-Arch can conduct end-to-end scientific research in the domain of architecture discovery, autonomously hypothesizing novel architectural concepts, implementing them as executable code, training and empirically validating their performance through rigorous experimentation and past experience. ASI-Arch conducted 1,773 autonomous experiments over 20,000 GPU hours, culminating in the discovery of 106 innovative, state-of-the-art (SOTA) linear attention architectures. Like AlphaGo's Move 37 that revealed unexpected strategic insights invisible to human players, our AI-discovered architectures demonstrate emergent design principles that systematically surpass human-designed baselines and illuminate previously unknown pathways for architectural innovation. Crucially, we establish the first empirical scaling law for scientific discovery itself--demonstrating that architectural breakthroughs can be scaled computationally, transforming research progress from a human-limited to a computation-scalable process. We provide comprehensive analysis of the emergent design patterns and autonomous research capabilities that enabled these breakthroughs, establishing a blueprint for self-accelerating AI systems.

架构发现AI for Science自进化AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。