首个自动评估大模型硬件编程能力的基准测试
PCEval: A Benchmark for Evaluating Physical Computing Capabilities of Large Language Models
- 构建全自动物理计算评估框架,覆盖逻辑与物理设计
- 13个主流模型测试显示代码生成强但电路布线差
- 适合硬件教育、AI辅助设计研究者使用
大型语言模型(LLMs)在软件开发等领域的表现引人注目,但在涉及硬件交互的物理计算场景中,其实际效果尚未充分探索。为此,我们提出 extsc{PCEval}(Physical Computing Evaluation),首个可自动评估 LLM 在逻辑与物理层面项目能力的基准测试,无需人工判断。该框架在不同复杂度下评估模型生成电路和兼容代码的能力。通过对13个领先模型的全面测试, extsc{PCEval} 提供了首个可复现且自动验证的实证评估,揭示了大模型在仿真环境中对基础硬件实现约束的推理能力。结果表明:尽管在代码生成和逻辑电路设计上表现良好,但在布线布局上存在显著缺陷,尤其在引脚连接管理与电路错误规避方面表现不佳。 extsc{PCEval} 深化了对AI在依赖硬件计算环境中的辅助作用的理解,并为开发更有效的物理计算教育工具奠定了基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, including software development, education, and technical assistance. Among these, software development is one of the key areas where LLMs are increasingly adopted. However, when hardware constraints are considered-for instance, in physical computing, where software must interact with and control physical hardware -their effectiveness has not been fully explored. To address this gap, we introduce \textsc{PCEval} (Physical Computing Evaluation), the first benchmark in physical computing that enables a fully automatic evaluation of the capabilities of LLM in both the logical and physical aspects of the projects, without requiring human assessment. Our evaluation framework assesses LLMs in generating circuits and producing compatible code across varying levels of project complexity. Through comprehensive testing of 13 leading models, \textsc{PCEval} provides the first reproducible and automatically validated empirical assessment of LLMs' ability to reason about fundamental hardware implementation constraints within a simulation environment. Our findings reveal that while LLMs perform well in code generation and logical circuit design, they struggle significantly with physical breadboard layout creation, particularly in managing proper pin connections and avoiding circuit errors. \textsc{PCEval} advances our understanding of AI assistance in hardware-dependent computing environments and establishes a foundation for developing more effective tools to support physical computing education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。