用多个大模型生成硬件代码,通过形式化验证筛选正确设计。
NoTB: Oracle-Free Triage of LLM-Generated RTL via Cross-Model Formal Consensus

- 多个独立训练的大模型生成RTL,用形式化方法比对一致性。
- 四模型一致时达94.7%准确率,三模型一致时87%准确率。
- 无需测试平台或黄金参考,适合早期硬件设计评估。
大型语言模型(LLMs)正被用于从自然语言规范生成寄存器传输级(RTL)设计,但早期功能正确性评估仍是根本挑战。现有无金标准方法依赖仿真一致性,受生成测试平台影响大;或使用大模型作为评判者,结果不一致。本文提出NoTB框架,通过跨模型形式化共识实现无金标准的早期筛选。该方法从多个独立训练的LLM家族生成RTL,并采用顺序等价检查(SEC)识别可证明等价的设计。我们发现,同一等价簇内模型族多样性带来校准的正确性信号,可在不依赖测试平台的情况下实现风险与覆盖的权衡。在78个CVDP RTL生成任务中,四族共识达到94.7%精度、27%覆盖率;三族共识达87%精度、33%覆盖率。这些性能点使设计师可在可信测试平台或黄金RTL出现前,灵活制定接受或延后决策规则。总体表明,形式化跨模型一致可为高置信度筛选提供可靠基础。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to generate register-transfer-level (RTL) designs from natural-language specifications. However, assessing functional correctness at early stages remains a fundamental challenge. Existing oracle-free approaches rely either on simulation-based agreement, which depends on LLM-generated testbenches that can fail or vary across models, or on LLM-as-a-judge heuristics, which produce inconsistent predictions. We introduce NoTB, an oracle-free triage framework that infers correctness from cross-model formal consensus. NoTB generates RTL implementations from multiple independently trained LLM families and applies Sequential Equivalence Checking (SEC) to identify designs that are provably equivalent. We show that the diversity of model families within an SEC-equivalent cluster induces a calibrated correctness signal, enabling risk-coverage tradeoffs without requiring testbenches. On 78 CVDP RTL-generation tasks, four-family formal consensus achieves 94.7% precision at 27% coverage; three-family consensus achieves 87% precision at 33% coverage. These operating points give designers a tunable accept/defer rule before a trusted testbench or golden RTL is available. Overall, NoTB demonstrates that formal cross-model agreement provides a reliable basis for high-confidence triage without model-dependent oracles
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。