AI代理在芯片设计全流程中表现参差,效率差距可达141倍。
Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows

- 用大模型代理执行从RTL到GDS的全流程芯片设计
- 同一进度下代理耗时与成本差异高达141倍
- 工具接口不匹配是物理设计失败主因,适合芯片自动化研究者
大语言模型代理正将电子设计自动化(EDA)从静态RTL生成扩展至长周期、工具交互式工作流。然而,通用编码代理即使具备领域特定技能,能否可靠完成涵盖综合、物理实现和工程变更优化的端到端RTL-to-GDS流程仍不明确。我们在PicoRV32设计上使用商用EDA工具,在两个时序目标下评估了三种代理架构和四种基础模型。通过端到端设计得分、阶段完成率及Token ROI(设计质量与运行时间、成本的关联度)进行评估。结果揭示三个关键教训:第一,领域技能提升代理对子任务的理解,但无法保证长流程可靠完成;第二,相似设计进展的代理间Token ROI最高相差141倍,显示运行时间和成本效率差异巨大;第三,低层工具接口不匹配是物理设计失败的主要原因,尤其当Tcl命令依赖工具版本或执行模式时。结论表明,鲁棒的智能体式EDA需更强模型,还需结构化工具接口、持续设计上下文、可控执行与流程级评估。
原文摘要 · Abstract (English)
Large language model (LLM) agents are extending electronic design automation (EDA) beyond static RTL generation toward long-horizon, tool-interactive workflows. Yet it remains unclear whether general-purpose coding agents, even with domain-specific EDA skills, can reliably execute an end-to-end RTL-to-GDS flow encompassing synthesis, physical implementation, and engineering change order (ECO) optimization. We evaluate AI agents on a PicoRV32 RTL-to-GDS flow using commercial EDA tools under two timing targets. Their performance is assessed using end-to-end design score, stage completion, and Token ROI, a cost-efficiency metric relating design quality to runtime and cost. Comparing three agent architectures and four foundation models, we derive three practical lessons. First, domain-specific skills improve agents' understanding of individual subtasks but do not ensure reliable completion of a long-horizon EDA flow. Second, agents that achieve similar design progress can still differ by up to 141 times in Token ROI, revealing substantial differences in runtime and cost efficiency. Third, low-level tool-interface mismatches are a major source of physical design failures, particularly when Tcl commands depend on the tool version or execution mode. These results suggest that robust Agentic EDA requires not only stronger models but also structured tool interfaces, persistent design context, controlled execution, and process-level evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。