arXiv:2602.08887cs.SEcs.AI2026-02

用大模型评估用户故事质量,专家认可其判断力但需更好融入工作流。

DeepQuali: Initial results of a study on the use of large language models for assessing the quality of user stories

  • 用GPT-4o结合质量模型自动评估用户故事质量
  • 专家与大模型在整体评分和解释上高度一致
  • 适合敏捷团队快速验证需求质量,但需改进工具集成

生成式人工智能(GAI),特别是大语言模型(LLM),在软件工程中主要用于编码任务。然而,需求工程——尤其是需求验证——对GAI的应用仍有限。当前使用GAI进行需求工作的重点集中在需求获取、转换和分类,而非质量评估。本文提出并评估了一种基于LLM(GPT-4o)的“DeepQuali”方法,用于在敏捷开发中评估和提升需求质量。我们在两家小型公司的项目中应用该方法,并将大模型的质量评估结果与专家判断进行对比。专家还参与了方案走查,提供反馈并评定对方法的接受度。结果显示,专家与大模型在整体评分及解释方面高度一致,但在细节评分上存在分歧,表明经验和专业性会影响判断。专家认为该方法有实用价值,但批评其缺乏工作流集成。大模型在支持工程师进行需求质量评估与改进方面具有潜力,明确使用质量模型并提供解释性反馈可提高接受度。

原文摘要 · Abstract (English)

Generative artificial intelligence (GAI), specifically large language models (LLMs), are increasingly used in software engineering, mainly for coding tasks. However, requirements engineering - particularly requirements validation - has seen limited application of GAI. The current focus of using GAI for requirements is on eliciting, transforming, and classifying requirements, not on quality assessment. We propose and evaluate the LLM-based (GPT-4o) approach "DeepQuali", for assessing and improving requirements quality in agile software development. We applied it to projects in two small companies, where we compared LLM-based quality assessments with expert judgments. Experts also participated in walkthroughs of the solution, provided feedback, and rated their acceptance of the approach. Experts largely agreed with the LLM's quality assessments, especially regarding overall ratings and explanations. However, they did not always agree with the other experts on detailed ratings, suggesting that expertise and experience may influence judgments. Experts recognized the usefulness of the approach but criticized the lack of integration into their workflow. LLMs show potential in supporting software engineers with the quality assessment and improvement of requirements. The explicit use of quality models and explanatory feedback increases acceptance.

需求工程大模型应用质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。