arXiv:2512.20823cs.ARcs.AI2025-12被引 2

构建可持续更新的硬件代码生成评测基准,推动大模型在真实电路设计中的应用。

NotSoTiny: A Large, Living Benchmark for RTL Code Generation

  • 基于真实硬件项目构建自动化评测流水线,避免数据污染。
  • 包含数百个结构复杂的设计,挑战远超以往基准。
  • 适合关注大模型在芯片设计中落地的科研与工程人员。

大语言模型在生成硬件描述语言(RTL)代码方面展现出早期潜力,但现有评估仍面临规模小、设计简单、验证不严及数据污染等问题。为此,本文提出 NotSoTiny 基准,用于评估大模型在生成结构丰富且上下文敏感的 RTL 代码方面的能力。该基准源自 Tiny Tapeout 社区的数百个真实硬件设计,通过自动化流程去除重复项、验证正确性,并按 Tiny Tapeout 发布节奏定期更新,有效防止数据污染。实验表明,NotSoTiny 的任务难度显著高于已有基准,充分揭示当前大模型在硬件设计中的局限性,同时为该技术的改进提供可靠指引。

原文摘要 · Abstract (English)

LLMs have shown early promise in generating RTL code, yet evaluating their capabilities in realistic setups remains a challenge. So far, RTL benchmarks have been limited in scale, skewed toward trivial designs, offering minimal verification rigor, and remaining vulnerable to data contamination. To overcome these limitations and to push the field forward, this paper introduces NotSoTiny, a benchmark that assesses LLM on the generation of structurally rich and context-aware RTL. Built from hundreds of actual hardware designs produced by the Tiny Tapeout community, our automated pipeline removes duplicates, verifies correctness and periodically incorporates new designs to mitigate contamination, matching Tiny Tapeout release schedule. Evaluation results show that NotSoTiny tasks are more challenging than prior benchmarks, emphasizing its effectiveness in overcoming current limitations of LLMs applied to hardware design, and in guiding the improvement of such promising technology.

RTL生成大模型硬件验证基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。