评测硬件智能体在预算内跨阶段设计闭合的能力,揭示其真实瓶颈。
CLOSER-Bench: Evaluating Budgeted Cross-Stage Design Closure for Hardware Agents
- 构建可控的跨阶段设计闭合评估流程,涵盖从规格到GDS全链路
- 三款智能体在局部AXI修复任务中完成率超90%,但验证闭合任务失败率超50%
- 适合关注硬件自动化、智能体决策能力的研究者与工程师
硬件工程对编码智能体构成长周期任务挑战:进展连续,工具反馈延迟且异构,后端失败可能需修改RTL而非调参。现有基准仅衡量RTL生成、代码库修复、验证、PPA演进或物理实现,因设计和评判标准不同,难以判断智能体在抽象边界上的表现。我们提出CLOSER-Bench,一个针对预算内跨阶段设计闭合的受控评估协议。针对同一设计与隐藏目标,该基准包含spec-to-RTL、RTL-to-GDS、spec-to-GDS三类任务,记录每次仿真、综合、时序分析及布局布线调用,度量最终质量、任意时刻进度、工具开销及跨阶段恢复能力。基准基于开源工具链:Verilator、Yosys、OpenROAD、KLayout、Sky130与Harbor智能体框架。十项预研任务覆盖RTL修复、基于变异的验证、覆盖率提升、PPA优化、设计空间探索、跨模型调试与安全检测,验证了可执行框架,并暴露显著的完成-闭合差距:三款智能体在局部AXI修复任务中成功率达90%以上,而匹配的验证闭合任务使顶尖智能体与两个此前成功的基线产生明显分离。进一步验证了完整的RTL-to-GDS流程,并构建宏级AXI/DMA流加速器用于阶段配对评估。结果表明,应将硬件闭合视为预算约束下的序列决策问题,而非独立的代码生成任务。
原文摘要 · Abstract (English)
Hardware engineering exposes coding agents to a form of long-horizon work that is difficult to capture with pass-at-k: progress is continuous, tool feedback is delayed and heterogeneous, and a backend failure may require revising RTL rather than tuning another physical-design parameter. Existing benchmarks measure RTL generation, repository repair, verification, PPA evolution, or physical implementation, but their different designs and oracles make it hard to determine where an agent succeeds or fails across abstraction boundaries. We introduce CLOSER-Bench, a controlled evaluation protocol for budgeted cross-stage design closure. For one design and one hidden objective, it pairs spec-to-RTL, RTL-to-GDS, and spec-to-GDS tasks, records every simulator, synthesis, STA, and place-and-route invocation, and measures final quality, anytime progress, tool cost, and cross-stage recovery. The benchmark is built on open-source Verilator, Yosys, OpenROAD, KLayout, Sky130, and the Harbor agent harness. A ten-task pilot spanning RTL repair, mutation-based verification, coverage, PPA optimization, design-space exploration, cross-model debugging, and security establishes the executable harness and exposes a sharp completion--closure gap: three agents solve a localized AXI repair task, while the matched verification-closure task separates a frontier agent from two otherwise successful baselines. We further validate a full RTL-to-GDS flow and construct a macro-based AXI/DMA streaming accelerator for the stage-paired evaluation. These results motivate treating hardware closure as a budgeted sequential decision problem rather than a collection of independent code generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。