arXiv:2605.15537cs.AI2026-05中稿 · DAC 2026被引 1

用智能代理自动修复和更新硬件生成基准,降低人工维护成本。

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision

论文配图:RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision
图 1 · 摘自论文原文
  • 引入智能体框架自动识别并修正错误的基准案例
  • 可检测并更新模型过拟合的测试用例,提升基准可靠性
  • 适合EDA研究者与芯片设计工具开发者使用

本文提出RTL-BenchMT,一种用于动态维护RTL生成基准的智能体框架。大语言模型辅助的自动化RTL生成是EDA领域的重要方向,但现有基准面临两大挑战:(1)基准中存在错误案例;(2)模型对基准产生过拟合。这两类问题仅靠人工难以有效解决。为系统性降低人工维护成本,我们设计了自动化智能体框架RTL-BenchMT,聚焦两项核心任务:(1)自动识别并修正错误的基准案例;(2)自动检测并更新过拟合案例。借助该框架,我们对错误与过拟合案例进行了深入分析,并构建了一个优化后的基准套件,将向社区开源。

原文摘要 · Abstract (English)

This paper introduces RTL-BenchMT, an agentic framework for dynamically maintaining RTL generation benchmarks. Large Language Models (LLMs) assisted automated RTL generation is one of the most important directions in EDA research. However, current RTL benchmarks face two critical challenges: (1) flawed cases in the benchmarks and (2) overfitting to the benchmarks. Both challenges are difficult to resolve purely by manual engineering effort. To address these issues and systematically reduce human maintenance costs, we propose an automated agentic framework, RTL-BenchMT. RTL-BenchMT focuses on two key applications: (1) automatically identifying and revising flawed benchmark cases and (2) automatically detecting and updating overfitting cases. With the assistance of RTL-BenchMT, we conduct a thorough, in-depth analysis of flawed and overfitting cases and produce a refined benchmark suite that will be open-sourced to the community.

EDA智能体基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。