Saarthi框架通过多智能体协作提升形式化验证效率,准确率提升70%。
Saarthi for AGI: Towards Domain-Specific General Intelligence for Formal Verification
- 引入规则书与语法约束,精准生成SystemVerilog断言
- 融合GraphRAG技术,使断言生成迭代次数减少50%
- 专为短时短上下文任务设计,适合验证工程师使用
Saarthi是一个基于多智能体协作的代理型AI框架,用于端到端形式化验证。尽管当前框架在从规范到覆盖率闭合的全流程中已实现约40%的成效,仍需改进以增强鲁棒性。本文提出两项关键优化:(1) 引入结构化规则书与SVA语法,提升SystemVerilog断言生成的准确性和可控性;(2) 集成GraphRAG等高级检索增强生成技术,使智能体可访问技术知识与最佳实践,实现输出的迭代优化。在针对NVIDIA CVDP基准的挑战性测试用例上进行评估,结果显示断言生成准确率提升70%,达成覆盖率闭合所需的迭代次数减少50%。
原文摘要 · Abstract (English)
Saarthi is an agentic AI framework that uses multi-agent collaboration to perform end-to-end formal verification. Even though the framework provides a complete flow from specification to coverage closure, with around 40% efficacy, there are several challenges that need to be addressed to make it more robust and reliable. Artificial General Intelligence (AGI) is still a distant goal, and current Large Language Model (LLM)-based agents are prone to hallucinations and making mistakes, especially when dealing with complex tasks such as formal verification. However, with the right enhancements and improvements, we believe that Saarthi can be a significant step towards achieving domain-specific general intelligence for formal verification. Especially for problems that require Short Term, Short Context (STSC) capabilities, such as formal verification, Saarthi can be a powerful tool to assist verification engineers in their work. In this paper, we present two key enhancements to the Saarthi framework: (1) a structured rulebook and specification grammar to improve the accuracy and controllability of SystemVerilog Assertion (SVA) generation, and (2) integration of advanced Retrieval Augmented Generation (RAG) techniques, such as GraphRAG, to provide agents with access to technical knowledge and best practices for iterative refinement and improvement of outputs. We also benchmark these enhancements for the overall Saarthi framework using challenging test cases from NVIDIA's CVDP benchmark targeting formal verification. Our benchmark results stand out with a 70% improvement in the accuracy of generated assertions, and a 50% reduction in the number of iterations required to achieve coverage closure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。