arXiv:2606.29520cs.SEcs.AI2026-06

首个评估大模型软件架构理解能力的开源基准

SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models

论文配图:SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models
图 1 · 摘自论文原文
  • 构建2154道专家标注多选题,覆盖8类架构知识与4种上下文长度
  • 11个模型零样本与五样本测试,准确率高但各领域表现差异大
  • 适合研究者和开发者评估大模型在系统设计中的专业推理能力

大型语言模型(LLMs)在软件开发全周期中日益作为助手使用,但其对软件架构的推理能力尚未被系统衡量。架构决策依赖于质量属性权衡、设计模式和系统级约束,而现有基准多聚焦语法或算法任务,未能覆盖这些内容。我们提出SAKE(Software Architectural Knowledge Evaluation),一个标准化且可复现的基准,用于评估LLMs的软件架构知识。SAKE包含2154道专家精心设计的多选题,每题4个选项,按8类架构主题和4种上下文长度分层。我们在零样本与五样本设置下评估了11个专有及开源权重模型。整体准确率较高,但不同类别间表现差异显著,暴露出专业实践中关键领域的能力短板。SAKE及其评估脚本和所有结果均开源,为社区提供追踪大模型架构推理能力的基准。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used as assistants across the software development lifecycle, yet their ability to reason about software architecture remains largely unmeasured. Architectural decision-making depends on quality attribute trade-offs, design patterns, and system-level constraints, none of which are exercised by benchmarks that target syntactic or algorithmic tasks. We introduce SAKE (Software Architectural Knowledge Evaluation), a standardized and reproducible benchmark for assessing software architectural knowledge in LLMs. SAKE comprises 2154 expert-curated multiple-choice questions, each with four options, stratified across eight architectural categories and four context-length levels. We evaluate 11 proprietary and open-weight models in zero-shot and five-shot settings. Overall accuracy is high, but performance varies markedly across categories, revealing competency gaps in areas central to professional practice. SAKE, its evaluation scripts, and all results are released as open source to give the community a baseline for tracking architectural reasoning in LLMs.

大模型评估软件架构开源基准LLM能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。