arXiv:2502.07980cs.LGcs.AI2025-02被引 20

首个面向模拟电路的LLM推理评测基准,揭示大模型在电路理解上的显著短板。

CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs

  • 构建510组电路问答对,覆盖多层级模拟电路知识
  • GPT-4o在数值答案上仅达48.04%准确率,单元测试通过率仅27.45%
  • 专设单元测试机制,凸显大模型在拓扑推理中的不足

大型语言模型(LLMs)在模拟电路设计中的作用尚未得到充分探索,而该领域正亟需超越传统优化方法的基于推理的解决方案。尽管相关性日益增强,但目前尚无评估LLMs电路推理能力的基准。为此,我们构建了包含510个问题-答案对的CIRCUIT数据集,涵盖多种模拟电路相关主题。在该数据集上表现最佳的模型GPT-4o,在最终数值答案上的准确率为48.04%。为评估模型鲁棒性,我们引入类似单元测试的评估机制,将问题分组进行测试;在此机制下,GPT-4o仅能通过27.45%的单元测试,表明最先进模型在理解电路——尤其是涉及电路拓扑结构时——仍面临挑战。该电路专用基准揭示了当前大模型在模拟集成电路设计应用中的局限性,为后续研究提供了重要参考。

原文摘要 · Abstract (English)

The role of Large Language Models (LLMs) has not been extensively explored in analog circuit design, which could benefit from a reasoning-based approach that transcends traditional optimization techniques. In particular, despite their growing relevance, there are no benchmarks to assess LLMs' reasoning capability about circuits. Therefore, we created the CIRCUIT dataset consisting of 510 question-answer pairs spanning various levels of analog-circuit-related subjects. The best-performing model on our dataset, GPT-4o, achieves 48.04% accuracy when evaluated on the final numerical answer. To evaluate the robustness of LLMs on our dataset, we introduced a unique feature that enables unit-test-like evaluation by grouping questions into unit tests. In this case, GPT-4o can only pass 27.45% of the unit tests, highlighting that the most advanced LLMs still struggle with understanding circuits, which requires multi-level reasoning, particularly when involving circuit topologies. This circuit-specific benchmark highlights LLMs' limitations, offering valuable insights for advancing their application in analog integrated circuit design.

电路推理LLM评测模拟电路基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。