arXiv:2605.06936cs.ARcs.AI2026-05被引 1

构建分层基准测试,评估大模型修复电路设计残余问题的能力

Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing

论文配图:Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing
图 1 · 摘自论文原文
  • 设计分层任务体系,覆盖规则修复与性能优化
  • 多模型测试显示复杂任务成功率不足四成
  • 视觉增强有效提升规则修复效果,推理能力是核心瓶颈

大模型正被用于电子设计自动化(EDA)的“最后一公里”:修复工具运行后遗留的签核设计规则检查(DRC)违规并收敛功耗-性能-面积(PPA)目标。现有EDA-LLM基准完全忽略DRC修复,且依赖单一工具链的扁平结构。本文提出PostEDA-Bench,一个包含145个任务的分层基准,涵盖DRC-Essential、DRC-Reasoning、PPA-Mono和PPA-Multi四类,支持机器可验证的EDA工具链评估。在八种商业与开源大模型、多种智能体架构下测试发现:模型对合成的DRC-Essential和单目标PPA-Mono任务表现尚可,但在更贴近实际的DRC-Reasoning任务中最佳成功率仅36.66%,在多目标PPA-Multi任务中最佳成功率仅为20.00%;视觉增强显著提升DRC-Bench表现;而权衡推理能力而非参数调优知识,是PPA-Multi任务的主要瓶颈。

原文摘要 · Abstract (English)

LLM-based agents are increasingly applied to the "last mile" of Electronic Design Automation (EDA): repairing residual sign-off Design Rule Check (DRC) violations and converging Power-Performance-Area (PPA) targets after tool runs. Existing EDA-LLM benchmarks, however, omit DRC fixing entirely and rely on flat hierarchies tied to a single toolchain. We introduce PostEDA-Bench, a hierarchical benchmark with 145 tasks across DRC-Essential, DRC-Reasoning, PPA-Mono, and PPA-Multi, supported by EDA toolchains with machine-checkable evaluation. Across eight commercial and open-source LLMs under multiple agent scaffolds, we find that agents handle synthetic DRC-Essential and single-objective PPA-Mono reasonably well but degrade sharply on the more practical DRC-Reasoning, where the best success rate is 36.66%, and PPA-Multi, where the best success rate is 20.00%; vision augmentation consistently enhances DRC-Bench; and trade-off reasoning, rather than knob knowledge, is the dominant PPA-Multi bottleneck.

EDA大模型电路设计基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。