arXiv:2506.14074cs.LGcs.AR2025-06被引 43

构建首个覆盖软硬件协同设计的大型Verilog基准数据集,推动AI在芯片设计中的落地应用。

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

  • 构建涵盖13类任务的783个真实芯片设计问题,含生成、验证与调试等场景。
  • 当前最先进模型代码生成准确率不足34%,尤其在复用与验证任务中表现差。
  • 专为评估大模型与智能体设计,适合芯片自动化与AI for Hardware研究者使用。

我们提出了综合Verilog设计问题(CVDP)基准数据集,旨在推动大语言模型和智能体在硬件设计与验证领域的研究。该数据集包含13个任务类别下的783个问题,覆盖RTL生成、验证、调试、规格对齐及技术问答,均由资深硬件工程师撰写。问题以非代理和代理两种格式提供。相较于以往工作,本基准引入更真实且更具挑战性的场景,当前最先进模型在代码生成上的pass@1指标最高仅为34%。特别是涉及RTL复用和验证的代理任务尤为困难。评估采用开源工具与模型评分框架,理解任务通过BLEU和基于LLM的评判进行评估。结果揭示了现有模型能力的巨大差距,凸显了向鲁棒性真实世界硬件设计自动化持续研究的必要性。

原文摘要 · Abstract (English)

We present the Comprehensive Verilog Design Problems (CVDP) benchmark, a new dataset and infrastructure to advance LLM and agent research in hardware design and verification. CVDP includes 783 problems across 13 task categories, covering RTL generation, verification, debugging, specification alignment, and technical Q&A authored by experienced hardware engineers. Problems are offered in both non-agentic and agentic formats. The benchmark introduces more realistic and challenging contexts than prior work, with state-of-the-art models achieving no more than 34% pass@1 on code generation. Agentic tasks$\unicode{x2013}$especially those involving RTL reuse and verification$\unicode{x2013}$are particularly difficult. Evaluation uses open-source tools and model scoring infrastructure, with comprehension tasks assessed via BLEU and LLM-based judging. CVDP reveals substantial gaps in current model capabilities, underscoring the need for continued research toward robust, real-world hardware design automation.

芯片设计大模型Verilog智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。