arXiv:2511.15767cs.LGcs.AI2025-11被引 2

用覆盖率引导的AI优化,自动生成高效硬件测试用例。

TB or Not TB: Coverage-Driven Direct Preference Optimization for Verilog Stimulus Generation

  • 用覆盖率反馈直接优化大模型生成测试代码。
  • 在基准上实现最高77.27%的代码覆盖率提升。
  • 适合芯片验证工程师快速生成高质量测试用例。

随着大语言模型(LLM)的快速发展,其在硬件设计与验证中的应用日益受到关注。其中,设计验证是最耗时且资源密集的环节,而为待测设计(DUT)生成有效测试激励既关键又费力。本文提出 { t TB or not TB} 框架,通过覆盖驱动的直接偏好优化(CD-DPO)微调 LLM 实现自动化测试激励生成。为支持基于偏好的训练,我们构建了 PairaNet 数据集,该数据集源自 PyraNet,通过仿真获得的覆盖率指标对高质量与低质量测试平台进行配对标注。所提出的 CD-DPO 方法将定量覆盖率反馈直接融入优化目标,引导模型生成能最大化验证覆盖率的激励。在 CVDP CID12 基准上的实验表明,{ t TB or not TB} 超过开源及商业基线,在代码覆盖率上最高提升达 77.27%,验证了覆盖驱动偏好优化在基于 LLM 的硬件验证中的有效性。

原文摘要 · Abstract (English)

With the rapid advancement of Large Language Models (LLMs), there is growing interest in applying them to hardware design and verification. Among these stages, design verification remains the most time-consuming and resource-intensive phase, where generating effective stimuli for the design under test (DUT) is both critical and labor-intensive. We present {\it TB or not TB}, a framework for automated stimulus generation using LLMs fine-tuned through Coverage-Driven Direct Preference Optimization (CD-DPO). To enable preference-based training, we introduce PairaNet, a dataset derived from PyraNet that pairs high- and low-quality testbenches labeled using simulation-derived coverage metrics. The proposed CD-DPO method integrates quantitative coverage feedback directly into the optimization objective, guiding the model toward generating stimuli that maximize verification coverage. Experiments on the CVDP CID12 benchmark show that {\it TB or not TB} outperforms both open-source and commercial baselines, achieving up to 77.27\% improvement in code coverage, demonstrating the effectiveness of Coverage-driven preference optimization for LLM-based hardware verification.

硬件验证大模型测试生成覆盖率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。