arXiv:2504.19384cs.SEcs.AI2025-04被引 7

用大模型提升需求工程中的定性分析效率,减少人工负担。

From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering

  • 采用上下文丰富的提示词,让大模型在演绎式标注中表现更优
  • GPT-4在演绎设置下与人工标注者一致性超0.7(Cohen's Kappa)
  • 生成的结构化标签可直接用于领域模型,支持需求可追溯

需求工程(RE)对复杂且受监管的软件项目至关重要。由于将利益相关者输入转化为一致设计存在挑战,定性数据分析(QDA)为处理自由文本数据提供了系统方法。然而传统QDA耗时且高度依赖人工。本文探讨了大语言模型(LLMs),包括GPT-4、Mistral和LLaMA-2,在需求工程中改进QDA任务的应用。研究评估了这些模型在归纳(零样本)与演绎(单样本、少样本)标注任务中的表现,发现GPT-4在演绎设置下与人类分析师达成显著一致性,Cohen's Kappa超过0.7,而零样本表现有限。详尽的上下文提示显著提升标注准确性和一致性,尤其在演绎场景中,GPT-4在多次运行中表现出高可靠性。结果表明,LLMs能有效支持需求工程中的QDA,降低人工成本同时保持标注质量。自动生成的结构化标签具备需求可追溯性,可直接作为领域模型中的类,促进系统化软件设计。

原文摘要 · Abstract (English)

Requirements Engineering (RE) is essential for developing complex and regulated software projects. Given the challenges in transforming stakeholder inputs into consistent software designs, Qualitative Data Analysis (QDA) provides a systematic approach to handling free-form data. However, traditional QDA methods are time-consuming and heavily reliant on manual effort. In this paper, we explore the use of Large Language Models (LLMs), including GPT-4, Mistral, and LLaMA-2, to improve QDA tasks in RE. Our study evaluates LLMs' performance in inductive (zero-shot) and deductive (one-shot, few-shot) annotation tasks, revealing that GPT-4 achieves substantial agreement with human analysts in deductive settings, with Cohen's Kappa scores exceeding 0.7, while zero-shot performance remains limited. Detailed, context-rich prompts significantly improve annotation accuracy and consistency, particularly in deductive scenarios, and GPT-4 demonstrates high reliability across repeated runs. These findings highlight the potential of LLMs to support QDA in RE by reducing manual effort while maintaining annotation quality. The structured labels automatically provide traceability of requirements and can be directly utilized as classes in domain models, facilitating systematic software design.

大模型需求工程定性分析自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。