QuArch数据集助语言模型理解计算机架构,提升研究效率。
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture
- 构建1500个经人工验证的问答对,覆盖处理器设计等关键领域。
- 闭源模型最高准确率84%,小开源模型达72%,内存系统仍存短板。
- 可用于训练和评估AI助手,适合计算机体系结构研究者使用。
我们提出QuArch,一个包含1500个经人工验证的问答对的数据集,用于评估和提升语言模型对计算机体系结构的理解能力。该数据集涵盖处理器设计、存储系统和性能优化等领域。分析显示显著性能差距:最佳闭源模型准确率达84%,顶级小型开源模型为72%。模型在存储系统、互连网络和基准测试方面表现较弱。使用QuArch进行微调可使小型模型准确率提升最多8%,为推动基于AI的计算机体系结构研究奠定基础。数据集与排行榜详见https://harvard-edge.github.io/QuArch/。
原文摘要 · Abstract (English)
We introduce QuArch, a dataset of 1500 human-validated question-answer pairs designed to evaluate and enhance language models' understanding of computer architecture. The dataset covers areas including processor design, memory systems, and performance optimization. Our analysis highlights a significant performance gap: the best closed-source model achieves 84% accuracy, while the top small open-source model reaches 72%. We observe notable struggles in memory systems, interconnection networks, and benchmarking. Fine-tuning with QuArch improves small model accuracy by up to 8%, establishing a foundation for advancing AI-driven computer architecture research. The dataset and leaderboard are at https://harvard-edge.github.io/QuArch/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。