arXiv:2607.00049cs.SEcs.AI2026-07

测试GPT-5在冲刺认证题上的表现,发现提示技巧能提升准确性。

Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study

论文配图:Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study
图 1 · 摘自论文原文
  • 用三种提示方法测试GPT-5回答冲刺认证题
  • 引用来源的提示准确率达89.1%,错误率最低
  • 适合敏捷学习者和备考者参考

大型语言模型(LLMs)在敏捷开发中的文档、指导和培训中应用日益广泛。随着从业者使用这些工具准备如专业冲刺大师(PSM)等认证,一个关键问题是LLMs能否可靠地理解冲刺框架——该框架规则明确,依据2020年《冲刺指南》。本文研究不同提示技术对LLM回答冲刺认证类问题事实准确性的影响。使用993个经验证的与PSM一致的问题集,采用零样本、思维链和带源引用三种提示方式让GPT-5作答。所有方法均达到超过85%的认证级准确率,其中基于引用的变体表现最佳(89.1%),错误率最低。正确答案集中在定义完成、事件、产品待办事项管理等明确主题,以及单选题;而多选题和更主观的领域如冲刺团队与产品价值则较不稳定。在至少一种提示失败的问题中(占16.2%),错误主要源于与《冲刺指南》不一致(28%)、超出其范围的内容(34%)及过时或偏见性解释(38%)。总体而言,提示技术带来温和但一致的提升,尤其减少误解和版本偏差,支持在敏捷学习与认证准备中更可靠地使用LLMs。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, coaching, and training. As practitioners adopt these tools to prepare for certifications such as Professional Scrum Master (PSM), a key question is whether LLMs can reliably reason about Scrum, a framework with normative, well-defined rules described in the Scrum Guide (2020). This paper examines how different prompt techniques affect the factual accuracy of LLM responses to Scrum certification-style questions. A dataset of 993 validated PSM-aligned questions was answered by GPT-5 using three techniques: zero-shot, chain-of-thought, and with-source citation. All prompts achieved certification-level accuracy above 85\%, with the citation-based variant performing best (89.1\%) and yielding the lowest error rate. Correct answers concentrated in well-defined topics, such as \emph{Definition of Done}, Events, and Product Backlog Management, and in single-answer multiple-choice items, while multi-select questions and more interpretive areas, such as Scrum Team and Product Value, were less stable. Among questions where at least one prompt failed (16.2\%), errors clustered into misalignment with the Scrum Guide (28\%), content outside its scope (34\%), and outdated or biased interpretations (38\%). Overall, prompt techniques produced modest but consistent improvements, particularly in reducing misinterpretation and version drift, supporting more reliable use of LLMs in Agile learning and certification preparation.

大模型敏捷开发提示工程认证备考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。