arXiv:2507.19525cs.LGcs.AI2025-07被引 6

首个聚焦电路设计的多模态模型评测基准,覆盖从原理到后端全流程。

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs

  • 构建3614个跨数字/模拟电路的问答对,按设计阶段与能力维度分类。
  • 实测显示大模型在后端设计与复杂计算上表现显著不足。
  • 适合从事EDA自动化与多模态模型研发的研究者使用。

多模态大语言模型(MLLMs)为电子设计自动化(EDA)带来了自动化与提升的潜力,但现有评测基准范围狭窄,难以全面评估其性能。为此,我们提出首个专用于电路设计的多模态评测基准MMCircuitEval,涵盖3614个精心筛选的问答对,覆盖数字与模拟电路在前端与后端设计等关键阶段,数据源自教材、技术题库、数据手册与真实文档,并经专家审核确保准确性和相关性。该基准按设计阶段、电路类型、测试能力(知识、理解、推理、计算)及难度分级,支持对模型能力与局限的细粒度分析。大量实验表明,现有大模型在后端设计与复杂计算任务中存在明显性能差距,凸显针对性训练数据与建模方法的必要性。MMCircuitEval为推动MLLM在EDA中的应用提供了基础资源,可集成至实际电路设计流程。代码与数据已开源:https://github.com/cure-lab/MMCircuitEval。

原文摘要 · Abstract (English)

The emergence of multimodal large language models (MLLMs) presents promising opportunities for automation and enhancement in Electronic Design Automation (EDA). However, comprehensively evaluating these models in circuit design remains challenging due to the narrow scope of existing benchmarks. To bridge this gap, we introduce MMCircuitEval, the first multimodal benchmark specifically designed to assess MLLM performance comprehensively across diverse EDA tasks. MMCircuitEval comprises 3614 meticulously curated question-answer (QA) pairs spanning digital and analog circuits across critical EDA stages - ranging from general knowledge and specifications to front-end and back-end design. Derived from textbooks, technical question banks, datasheets, and real-world documentation, each QA pair undergoes rigorous expert review for accuracy and relevance. Our benchmark uniquely categorizes questions by design stage, circuit type, tested abilities (knowledge, comprehension, reasoning, computation), and difficulty level, enabling detailed analysis of model capabilities and limitations. Extensive evaluations reveal significant performance gaps among existing LLMs, particularly in back-end design and complex computations, highlighting the critical need for targeted training datasets and modeling approaches. MMCircuitEval provides a foundational resource for advancing MLLMs in EDA, facilitating their integration into real-world circuit design workflows. Our benchmark is available at https://github.com/cure-lab/MMCircuitEval.

电路设计多模态评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。