arXiv:2507.21532cs.CLcs.AI2025-07中稿 · AIRE 2025被引 1

复现并扩展了用户需求分类的深度学习研究,验证了模型可复现性与泛化能力。

Automatic Classification of User Requirements from Online Feedback -- A Replication Study

  • 复现原始代码,评估不同模型在用户反馈中的需求分类表现。
  • BERT和ELMo在外部数据集上表现良好,GPT-4o性能接近传统模型。
  • 提供完整复现包与研究身份证,推动可复现研究生态建设。

自然语言处理技术被广泛应用于需求工程领域,支持分类与歧义检测等任务。尽管需求工程根植于实证研究,但对NLP用于需求工程(NLP4RE)的研究复现关注不足。随着NLP快速发展,机器辅助工作流带来新机遇。本研究复现并扩展了一项先前的NLP4RE研究,该研究评估了深度学习模型在小数据环境下从用户评论中分类需求的效果。我们使用公开源码重现了原始结果,增强了原研究的外部有效性。进一步在外部数据集上评估模型性能,并与GPT-4o零样本分类器进行对比。同时,为原研究准备了复现研究身份证,以评估其复现准备度。结果显示,不同模型复现程度各异,朴素贝叶斯完全可复现,而BERT及其他模型表现不一。基线模型如BERT和ELMo在外数据集上表现出良好泛化能力,GPT-4o性能与传统机器学习模型相当。评估确认原研究具备复现准备度;若缺少环境配置文件,准备度将受影响。本文补充了缺失信息,并提供本研究的复现身份证,以促进后续研究复现。

原文摘要 · Abstract (English)

Natural language processing (NLP) techniques have been widely applied in the requirements engineering (RE) field to support tasks such as classification and ambiguity detection. Although RE research is rooted in empirical investigation, it has paid limited attention to replicating NLP for RE (NLP4RE) studies. The rapidly advancing realm of NLP is creating new opportunities for efficient, machine-assisted workflows, which can bring new perspectives and results to the forefront. Thus, we replicate and extend a previous NLP4RE study (baseline), "Classifying User Requirements from Online Feedback in Small Dataset Environments using Deep Learning", which evaluated different deep learning models for requirement classification from user reviews. We reproduced the original results using publicly released source code, thereby helping to strengthen the external validity of the baseline study. We then extended the setup by evaluating model performance on an external dataset and comparing results to a GPT-4o zero-shot classifier. Furthermore, we prepared the replication study ID-card for the baseline study, important for evaluating replication readiness. Results showed diverse reproducibility levels across different models, with Naive Bayes demonstrating perfect reproducibility. In contrast, BERT and other models showed mixed results. Our findings revealed that baseline deep learning models, BERT and ELMo, exhibited good generalization capabilities on an external dataset, and GPT-4o showed performance comparable to traditional baseline machine learning models. Additionally, our assessment confirmed the baseline study's replication readiness; however missing environment setup files would have further enhanced readiness. We include this missing information in our replication package and provide the replication study ID-card for our study to further encourage and support the replication of our study.

需求工程复现研究NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。