arXiv:2504.16768cs.CLcs.AI2025-04被引 11

测试生成式大模型在需求分类任务中的表现,发现提示设计比模型类型更重要。

How Effective are Generative Large Language Models in Performing Requirements Classification?

  • 用三种生成式大模型测试分类效果,重点研究提示工程影响
  • 在三个数据集上完成400+实验,验证提示设计对性能影响显著
  • 适合关注LLM实际应用落地的软件工程研究人员

近年来,基于Transformer的大型语言模型(LLM)革新了自然语言处理,生成式模型为需要上下文感知文本生成的任务开辟了新可能。需求工程(RE)领域也涌现出大量对LLM的应用探索,涵盖追溯链接检测、合规性检查等任务。需求分类是RE中的常见任务。尽管非生成式模型如BERT已成功应用于该任务,但生成式模型的研究仍有限。这引发关键问题:具备上下文生成能力的生成式模型在需求分类中表现如何?本研究系统评估了Bloom、Gemma和Llama三款生成式模型在二分类与多分类任务中的表现。实验覆盖三个常用数据集(PROMISE NFR、Functional-Quality、SecReq),共开展400余次实验。结果表明,提示设计与模型架构普遍重要,而数据集差异的影响则因任务复杂度而异。该发现可指导未来模型开发与部署策略,强调优化提示结构并匹配任务特性的模型架构以提升性能。

原文摘要 · Abstract (English)

In recent years, transformer-based large language models (LLMs) have revolutionised natural language processing (NLP), with generative models opening new possibilities for tasks that require context-aware text generation. Requirements engineering (RE) has also seen a surge in the experimentation of LLMs for different tasks, including trace-link detection, regulatory compliance, and others. Requirements classification is a common task in RE. While non-generative LLMs like BERT have been successfully applied to this task, there has been limited exploration of generative LLMs. This gap raises an important question: how well can generative LLMs, which produce context-aware outputs, perform in requirements classification? In this study, we explore the effectiveness of three generative LLMs-Bloom, Gemma, and Llama-in performing both binary and multi-class requirements classification. We design an extensive experimental study involving over 400 experiments across three widely used datasets (PROMISE NFR, Functional-Quality, and SecReq). Our study concludes that while factors like prompt design and LLM architecture are universally important, others-such as dataset variations-have a more situational impact, depending on the complexity of the classification task. This insight can guide future model development and deployment strategies, focusing on optimising prompt structures and aligning model architectures with task-specific needs for improved performance.

需求工程大模型分类任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。