arXiv:2510.22087cs.ARcs.AI2025-10被引 3

首个专用于评估大模型计算机体系结构推理能力的基准测试

QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture

  • 构建2671个专家验证的问答对,覆盖处理器设计等多领域
  • 顶尖模型在复杂问题上准确率仅34%至73%,体现推理短板
  • 微调后可提升内存层次结构设计效率,最高省1.99倍面积

计算机体系结构作为高层软件抽象与底层硬件实现之间的桥梁,目前仍缺失于主流大语言模型(LLM)评估体系中。为此,我们提出QuArch(发音为'quark'),首个专为评估和推动大模型在计算机体系结构领域知识与推理能力而设计的基准测试。QuArch v1.0包含2,671个由专家验证的问答对,涵盖处理器设计、存储系统及互连网络等多个方面。评估发现,尽管前沿模型具备特定领域知识,但在需要高阶思维的任务中表现不佳:在分析、设计和实现类问题上,模型准确率在34%到73%之间波动,暴露出显著的体系结构推理差距。此外,通过微调训练,发现基于QuArch的模型可在真实内存层次结构设计任务中实现更优解,最多提升1.99倍面积效率,且整体可行解比例提高40%。该基准为构建和衡量大模型在计算系统创新中的能力提供了坚实基础。完整数据集与排行榜已公开:https://quarch.ai/

原文摘要 · Abstract (English)

The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) evaluations. To this end, we present QuArch (pronounced 'quark'), the first benchmark designed to facilitate the development and evaluation of LLM knowledge and reasoning capabilities specifically in computer architecture. QuArch v1.0 provides a comprehensive collection of 2,671 expert-validated question-answer (QA) pairs covering various aspects of computer architecture, including processor design, memory systems, and interconnection networks. Our evaluation reveals that while frontier models possess domain-specific knowledge, they struggle with skills that require higher-order thinking in computer architecture. Frontier model accuracies vary widely (from 34% to 73%) on these advanced questions, highlighting persistent gaps in architectural reasoning across analysis, design, and implementation QAs. Furthermore, via fine-tuning we find that QuArch can translate to improved performance on a realistic memory hierarchy design task, resulting in up to 1.99x more area-efficient solutions and up to 40% more viable solutions overall. By holistically assessing fundamental skills, QuArch provides a foundation for building and measuring LLM capabilities that can accelerate innovation in computing systems. The QuArch benchmark and leaderboard are publicly available at: https://quarch.ai/.

大模型评测体系结构推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。