提出首个结构验证的电路网表评测基准,揭示大模型在电路设计中的可靠性瓶颈。
NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation

- 构建结构验证的网表任务基准,涵盖参数识别、连接修改等24类操作。
- 局部修改准确率达96%~100%,设备添加仅41%~83%,长序列编辑性能骤降。
- 适合关注AI辅助电路设计可靠性的研究者与工程师使用。
大型语言模型(LLMs)在电路设计流程中应用日益广泛,但其在面向仿真器的SPICE网表识别与操作中的可靠性仍不明确,且常被高阶设计推理掩盖。尽管网表为文本形式,却通过拓扑与参数编码结构化电路对象。我们提出 extbf{NetlistBench},一个结构验证的SPICE网表识别与操作基准。该基准包含2,342个案例,覆盖24类任务,包括参数与连接识别与修改、层级操作、等价性判断及长时序复合编辑。模型输出由确定性结构感知的验证器评估。六种非思考型LLM在不同操作复杂度下表现差异显著:简单局部修改准确率达96%–100%,设备添加降至41%–83%,等价性判断为49%–90%。启用推理显著提升弱模型表现,但无法消除结构保持失败,且随着编辑跨度增加,性能急剧下降。NetlistBench将网表可靠性确认为可信LLM电路自动化的关键瓶颈。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist recognition and manipulation remains poorly understood and is rarely separated from high-level design reasoning. Although netlists are textual, they encode structured circuit objects through topology and parameters. We present \textbf{NetlistBench}, a structure-verified benchmark for SPICE netlist recognition and manipulation. NetlistBench contains 2,342 cases across 24 task families, covering parameter and connectivity recognition and edits, hierarchical operations, equivalence judgment, and long-horizon compound editing. Model outputs are evaluated by a deterministic structure-aware oracle. Across six non-thinking LLMs, performance varies substantially with operation-level structural complexity. Simple local edits reach $96\%$--$100\%$ accuracy, while device addition drops to $41\%$--$83\%$ and equivalence judgment to $49\%$--$90\%$. Enabling reasoning substantially improves weaker models but does not eliminate structure-preservation failures, with performance still degrading sharply as the edit horizon increases. NetlistBench identifies netlist reliability as a distinct bottleneck for trustworthy LLM-based circuit design automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。