构建首个全面评估多模态大模型在模拟电路能力的基准测试
AMSbench: A Comprehensive Benchmark for Evaluating MLLM Capabilities in AMS Circuits
- 设计涵盖电路图理解、分析与设计的8000道题目
- 发现现有模型在复杂多模态推理中表现明显不足
- 适合从事模拟电路自动化设计的研究者参考
模拟/混合信号(AMS)电路在集成电路产业中至关重要,但其自动化设计长期面临挑战。尽管多模态大语言模型(MLLMs)展现出支持AMS电路分析与设计的潜力,当前研究多局限于单一任务评估,缺乏系统性基准。为此,我们提出AMSbench,一个涵盖电路图感知、电路分析与电路设计等关键任务的评测基准,包含约8000个不同难度级别的测试题,评估了包括Qwen 2.5-VL和Gemini 2.5 Pro在内的八种主流模型。评估结果揭示当前MLLM在复杂多模态推理与高级电路设计任务中存在显著局限,凸显提升模型对电路领域知识理解的必要性,以缩小与人类专家的性能差距,推动实现全自动化AMS电路设计流程。数据已公开。
原文摘要 · Abstract (English)
Analog/Mixed-Signal (AMS) circuits play a critical role in the integrated circuit (IC) industry. However, automating Analog/Mixed-Signal (AMS) circuit design has remained a longstanding challenge due to its difficulty and complexity. Although recent advances in Multi-modal Large Language Models (MLLMs) offer promising potential for supporting AMS circuit analysis and design, current research typically evaluates MLLMs on isolated tasks within the domain, lacking a comprehensive benchmark that systematically assesses model capabilities across diverse AMS-related challenges. To address this gap, we introduce AMSbench, a benchmark suite designed to evaluate MLLM performance across critical tasks including circuit schematic perception, circuit analysis, and circuit design. AMSbench comprises approximately 8000 test questions spanning multiple difficulty levels and assesses eight prominent models, encompassing both open-source and proprietary solutions such as Qwen 2.5-VL and Gemini 2.5 Pro. Our evaluation highlights significant limitations in current MLLMs, particularly in complex multi-modal reasoning and sophisticated circuit design tasks. These results underscore the necessity of advancing MLLMs' understanding and effective application of circuit-specific knowledge, thereby narrowing the existing performance gap relative to human expertise and moving toward fully automated AMS circuit design workflows. Our data is released at this URL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。