arXiv:2608.26418cs.ARcs.AI2026-08

AI自动生成从设计到部署的完整推理芯片,两周完成。

Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI

  • AI驱动软硬件协同设计,全程自动优化验证。
  • 红木芯片实测性能功耗比提升3.4倍,支持千亿参数模型。
  • 可快速迭代更新,适合追求极致效率的AI硬件研发。

现代AI工作负载与运行它们的硬件演进节奏不同:架构定义滞后于大规模硅片数年,而目标工作负载却在数月内变化。因此,设计决策在高度不确定下做出,且需付出双重代价——一次是为应对不确定性增加通用性,另一次是新工作负载难以适配已冻结的硬件。随着摩尔定律停滞,专用化成为提升能效比的唯一主要途径,要求设计周期与工作负载同步。我们提出一个端到端的AI系统,将软件到硅片的链条整合为单一优化循环,使硬件与软件在统一目标下协同设计并验证。其首个成果是红木(Redwood),一款面向物理智能的单批处理、低功耗、超低延迟推理的前沿加速器。仅凭两名人类架构师提供的高层规格,该系统在两周内自主生成了性能模型、RTL设计、UVM环境、形式化证明、固件和内核,全程无人干预。每个模块均通过商业EDA工具、自研形式化引擎及硬件在环验证达到95%覆盖率。规格变更可在48小时内重新验证并部署至硬件。红木纳米版(Nano)为超低功耗FPGA版本,可运行如Llama和Qwen等千亿参数模型。若映射至三星8纳米工艺(类Jetson Orin Nano),红木相比实测基准的Jetson,在相同模型上实现1.75倍吞吐量,功耗降低1.9倍,能效比提升3.4倍。在红木上运行的Qwen还协助设计下一代红木,迈出递归自我改进的第一步。据我们所知,这是首个由AI系统端到端设计、生产可用并运行现代大模型的加速器。

原文摘要 · Abstract (English)

Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months. Design decisions are therefore committed under deep uncertainty and paid for twice, once in the generality added as a hedge, and again when new workloads map poorly onto frozen silicon. As Moore's Law stagnates, specialization is the main remaining source of performance-per-watt and demands a design cycle that runs at the cadence of the workloads. We present an end-to-end AI system that collapses the software-to-silicon stack into a single optimization loop, where hardware and software are co-designed and verified under one objective. Its first demonstration is Redwood, a frontier AI accelerator built for single-batch, low-power, ultra-low-latency inference for physical AI. From a high-level specification by two human architects, the system autonomously generated the performance model, RTL design, UVM environments, formal proofs, firmware, and kernels in under two weeks with no human intervention below the specification. Every block reached 95% coverage via commercial EDA tools, our proprietary formal engine, and hardware-in-the-loop validation. Specification changes were reverified and redeployed to hardware in under 48 hours. Redwood Nano, its ultra-low-power FPGA variant, runs multi-billion-parameter models like Llama and Qwen. Projected onto Samsung 8 nm, the Jetson Orin Nano's process class, Redwood delivers 1.75x the throughput at 1.9x lower power, a 3.4x performance-per-watt gain against a measured Jetson baseline on the same models. Qwen running on Redwood also helped design next-generation Redwood, an early step toward recursive self-improvement. To our knowledge, this is the first production-worthy AI accelerator designed end-to-end by an AI system and running a modern AI model.

AI芯片自动生成能效比硬件验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。