对比三款大模型在冲刺认证题上的表现,发现谷歌模型最准,但不同题型和知识点误差差异明显。
Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns

- 用三种提示策略测试三款大模型在993道冲刺认证题上的表现
- 谷歌模型准确率最高,多选与真假题错误较多,价值观类题目更易出错
- 错误具系统性:误读定义、过度泛化或混淆市场说法与严格规范
大型语言模型(LLMs)在考试与认证类问答任务中应用日益广泛,其获取、理解并应用领域知识的能力可被系统评估。在软件工程领域,此类任务尤其关键,因问题依赖于对规范定义、角色、工件和规则的严格遵循。本文评估了三款主流LLM——GPT-5 mini、Gemini 3 Flash与DeepSeek Chat 3.2——在993道符合专业冲刺大师Ⅰ级(PSM I)标准的冲刺认证题上的表现。采用零样本、思维链和源文本支撑三种提示策略,并通过重复执行评估模型内部稳定性。结果表明,各模型间性能差异显著:Gemini 3 Flash准确率最高,其次为GPT-5 mini,DeepSeek Chat 3.2最低;所有条件下模型内部变异度均较低。按题型看,单选多选题准确率最高,而多选与判断题更易出错。按主题划分,规范明确的领域如工件、经验主义与产品价值表现更稳,而冲刺价值观、自组织团队及利益相关者与客户等主题表现较弱。定性分析显示,错误具有系统性,包括过度泛化、措辞受限、复合干扰项以及市场通用理解与严格冲刺定义之间的冲突。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used in exam- and certification-style question answering tasks, where their ability to retrieve, interpret, and apply domain-specific knowledge can be systematically assessed. In Software Engineering, such settings are particularly relevant when questions depend on strict adherence to normative definitions, roles, artifacts, and rules. This paper evaluates the performance of three contemporary LLMs, \textit{GPT-5 mini}, \textit{Gemini 3 Flash}, and \textit{DeepSeek Chat 3.2}, in answering 993 Scrum certification-style questions aligned with the Professional Scrum Master I (PSM I) assessment format. We evaluated the models under three prompting strategies (\textit{zero-shot}, \textit{chain-of-thought}, and \textit{source-grounded}), with repeated executions to assess intra-model stability. We also analyzed performance across Scrum topics and question formats, complemented by a qualitative analysis of recurring error patterns in incorrect answers. Results revealed clear differences among models, with Gemini 3 Flash achieving the highest accuracy, followed by GPT-5 mini and DeepSeek Chat 3.2, while intra-model variability remained low across all conditions. By question format, the models achieved the highest accuracy on single-answer multiple-choice items, whereas multi-select and True/False questions were more error-prone. By topic, performance was more consistent in normatively explicit areas such as Artifacts, Empiricism, and Product Value, but more fragile in Scrum Values, Self-Managing Teams, and Stakeholders \& Customers. The qualitative analysis showed that errors were systematic rather than random, involving overgeneralization, restrictive wording, compound distractors, and conflicts between common market interpretations and strict Scrum definitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。