arXiv:2604.05755cs.SEcs.AI2026-04

首个评估大模型云架构理解能力的基准,揭示不同评测方式对结果的关键影响。

CAKE: Cloud Architecture Knowledge Evaluation of Large Language Models

  • 构建覆盖四认知层级与五云原生主题的188题基准CAKE
  • 3B以上参数模型在选择题准确率达99.2%,自由回答得分持续提升
  • 评测形式决定认知能力呈现:选择题有上限,开放题能区分模型差异

当前大型语言模型(LLMs)作为软件架构协作者,在云原生架构理解方面尚无可靠评估标准。为此,本文提出基准CAKE,包含188道专家验证题目,覆盖布卢姆修订版分类学中的四个认知层次(回忆、分析、设计、实现)及五个云原生主题。在22个模型配置(0.5B–70B参数)上进行评估,采用三轮多数投票处理多选题(MCQs),并使用大模型作为评判者评分自由回答(FR)。结果显示:第一,多选题准确率在3B以上参数时趋于饱和,最佳模型达99.2%;第二,自由回答得分随参数规模持续提升;第三,两种评测形式反映不同知识维度,多选题接近天花板而自由回答仍可区分模型;第四,推理增强(+think)提升自由回答质量,工具增强(+tool)却导致小模型性能下降。表明评测方式从根本上影响对模型架构知识的衡量。

原文摘要 · Abstract (English)

In today's software architecture, large language models (LLMs) serve as software architecture co-pilots. However, no benchmark currently exists to evaluate large language models' actual understanding of cloud-native software architecture. For this reason we present a benchmark called CAKE, which consists of 188 expert-validated questions covering four cognitive levels of Bloom's revised taxonomy -- recall, analyze, design, and implement -- and five cloud-native topics. Evaluation is conducted on 22 model configurations (0.5B--70B parameters) across four LLM families, using three-run majority voting for multiple-choice questions (MCQs) and LLM-as-a-judge scoring for free-responses (FR). Based on this evaluation, four notable findings were identified. First, MCQ accuracy plateaus above 3B parameters, with the best model reaching 99.2\%. Second, free-response scores scale steadily across all cognitive levels. Third, the two formats capture different facets of knowledge, as the MCQ accuracy approaches a ceiling while free-responses continue to differentiate models. Finally, reasoning augmentation (+think) improves free-response quality, while tool augmentation (+tool) degrades performance for small models. These results suggest that the evaluation format fundamentally shapes how we measure architectural knowledge in LLMs.

大模型评估云原生认知测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。