arXiv:2507.16200cs.LGcs.AR2025-07被引 19

首个面向真实IP设计的Verilog生成评测基准,揭示现有模型能力严重不足。

RealBench: Benchmarking Verilog Generation Models with Real-World IP Designs

  • 基于真实开源IP设计构建复杂、结构化数据集
  • 最先进模型在模块级仅13.3%通过率,系统级为0%
  • 适合硬件自动化、LLM应用研究者评估生成能力

利用大语言模型(LLMs)自动生成Verilog代码在硬件设计自动化领域备受关注。然而,现有评测基准因设计过于简单、规格不完整及验证环境不严谨,难以复现真实设计流程。为此,我们提出RealBench,首个专注于真实世界IP级Verilog生成任务的评测基准。它包含复杂、结构化的开源真实IP设计,支持多模态且格式化的设计规格,并配备严格的验证环境,包括100%行覆盖率测试平台和形式化检查器。该基准支持模块级与系统级任务,可全面评估LLM能力。对多种LLMs和智能体的评估显示,即使表现最佳的o1-preview模型,在模块级任务中也仅达13.3%的pass@1,系统级任务则为0%,凸显未来需更强的Verilog生成模型。基准已开源:https://github.com/IPRC-DIP/RealBench。

原文摘要 · Abstract (English)

The automatic generation of Verilog code using Large Language Models (LLMs) has garnered significant interest in hardware design automation. However, existing benchmarks for evaluating LLMs in Verilog generation fall short in replicating real-world design workflows due to their designs' simplicity, inadequate design specifications, and less rigorous verification environments. To address these limitations, we present RealBench, the first benchmark aiming at real-world IP-level Verilog generation tasks. RealBench features complex, structured, real-world open-source IP designs, multi-modal and formatted design specifications, and rigorous verification environments, including 100% line coverage testbenches and a formal checker. It supports both module-level and system-level tasks, enabling comprehensive assessments of LLM capabilities. Evaluations on various LLMs and agents reveal that even one of the best-performing LLMs, o1-preview, achieves only a 13.3% pass@1 on module-level tasks and 0% on system-level tasks, highlighting the need for stronger Verilog generation models in the future. The benchmark is open-sourced at https://github.com/IPRC-DIP/RealBench.

Verilog生成硬件自动化LLM评测真实设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。