arXiv:2510.21443cs.SEcs.AI2025-10被引 4

小模型在需求分类上表现接近大模型,更省资源且安全。

Does Model Size Matter? A Comparison of Small and Large Language Models for Requirements Classification

  • 对比8个模型,用三个数据集测试分类效果。
  • 小模型平均F1比大模型低2%,但无统计显著差异。
  • 小模型适合注重隐私和本地部署的工程场景。

大型语言模型(LLMs)在需求工程(RE)的自然语言处理任务中表现优异,但存在计算成本高、数据泄露风险及依赖外部服务等问题。小型语言模型(SLMs)则提供轻量、可本地部署的替代方案。本研究初步比较了3个LLM与5个SLM在需求分类任务上的表现,使用PROMISE、PROMISE Reclass和SecReq数据集。结果表明,尽管LLM平均F1分数比SLM高2%,但差异不具统计显著性;SLMs在所有数据集上几乎达到LLM性能,且在PROMISE Reclass数据集上召回率甚至更高,而模型规模最多缩小300倍。研究还发现,数据集特性对性能的影响大于模型大小。本研究为SLMs作为需求分类的有效替代方案提供了实证支持,其优势在于隐私保护、成本更低及本地部署可行性。

原文摘要 · Abstract (English)

[Context and motivation] Large language models (LLMs) show notable results in natural language processing (NLP) tasks for requirements engineering (RE). However, their use is compromised by high computational cost, data sharing risks, and dependence on external services. In contrast, small language models (SLMs) offer a lightweight, locally deployable alternative. [Question/problem] It remains unclear how well SLMs perform compared to LLMs in RE tasks in terms of accuracy. [Results] Our preliminary study compares eight models, including three LLMs and five SLMs, on requirements classification tasks using the PROMISE, PROMISE Reclass, and SecReq datasets. Our results show that although LLMs achieve an average F1 score of 2% higher than SLMs, this difference is not statistically significant. SLMs almost reach LLMs performance across all datasets and even outperform them in recall on the PROMISE Reclass dataset, despite being up to 300 times smaller. We also found that dataset characteristics play a more significant role in performance than model size. [Contribution] Our study contributes with evidence that SLMs are a valid alternative to LLMs for requirements classification, offering advantages in privacy, cost, and local deployability.

需求工程小模型分类任务本地部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。