arXiv:2503.15248cs.SEcs.AI2025-03被引 9

用大模型自动从功能需求生成质量类非功能需求,提升软件质量

Automated Non-Functional Requirements Generation in Software Engineering with Large Language Models: A Comparative Study

  • 基于大模型和自定义提示词,从功能需求推导出质量属性
  • 生成1593条非功能需求,专家评估有效性达4.63分(满分5分)
  • 适合软件需求工程师和AI辅助开发研究者参考

早期忽视非功能需求(NFRs)会引发严重问题。尽管其重要性突出,但常被忽略或难以识别。为此,我们构建了一个基于大语言模型(LLMs)的框架,通过自定义提示技术,在Deno管道中从功能需求(FRs)推导出以质量为导向的NFRs,实现系统化集成。针对生成质量,采用横向评估,涵盖三方面:NFR有效性、质量属性适用性、分类精确度。由10位平均13年经验的行业专家评估部分结果,显示生成的NFR与专家判断高度一致,中位有效性与适用性评分均为5.0(均值分别为4.63和4.59)。在分类任务中,80.4%的属性匹配专家选择,8.3%近似匹配,11.3%不匹配。对八种模型的对比分析表明性能差异显著,gemini-1.5-pro在属性准确率上最优,llama-3.3-70B则在有效性和适用性上表现更佳。研究验证了使用大模型自动化生成非功能需求的可行性,为智能需求工程发展奠定基础。

原文摘要 · Abstract (English)

Neglecting non-functional requirements (NFRs) early in software development can lead to critical challenges. Despite their importance, NFRs are often overlooked or difficult to identify, impacting software quality. To support requirements engineers in eliciting NFRs, we developed a framework that leverages Large Language Models (LLMs) to derive quality-driven NFRs from functional requirements (FRs). Using a custom prompting technique within a Deno-based pipeline, the system identifies relevant quality attributes for each functional requirement and generates corresponding NFRs, aiding systematic integration. A crucial aspect is evaluating the quality and suitability of these generated requirements. Can LLMs produce high-quality NFR suggestions? Using 34 functional requirements - selected as a representative subset of 3,964 FRs-the LLMs inferred applicable attributes based on the ISO/IEC 25010:2023 standard, generating 1,593 NFRs. A horizontal evaluation covered three dimensions: NFR validity, applicability of quality attributes, and classification precision. Ten industry software quality evaluators, averaging 13 years of experience, assessed a subset for relevance and quality. The evaluation showed strong alignment between LLM-generated NFRs and expert assessments, with median validity and applicability scores of 5.0 (means: 4.63 and 4.59, respectively) on a 1-5 scale. In the classification task, 80.4% of LLM-assigned attributes matched expert choices, with 8.3% near misses and 11.3% mismatches. A comparative analysis of eight LLMs highlighted variations in performance, with gemini-1.5-pro exhibiting the highest attribute accuracy, while llama-3.3-70B achieved higher validity and applicability scores. These findings provide insights into the feasibility of using LLMs for automated NFR generation and lay the foundation for further exploration of AI-assisted requirements engineering.

需求工程大模型非功能需求AI辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。