arXiv:2501.06491cs.SEcs.AI2025-01被引 13

用SMOTE-Tomek提升需求分类准确率

Improving Requirements Classification with SMOTE-Tomek Preprocessing

  • 结合SMOTE-Tomek与分层交叉验证处理数据不平衡问题
  • 逻辑回归准确率达76.16%,较基线59.85%显著提升
  • 适合关注软件需求工程与机器学习应用的研究者

本研究聚焦需求工程领域,采用SMOTE-Tomek预处理技术,并结合分层K折交叉验证,解决PROMISE数据集中的类别不平衡问题。该数据集包含969条已分类的需求,分为功能型与非功能型两类。所提方法在保持验证折完整性的同时,有效增强了少数类的表示能力,显著提升了分类准确率。逻辑回归模型达到76.16%的准确率,明显优于基线59.85%。结果表明,机器学习模型可作为可扩展且可解释的解决方案,在需求分类中具有高效性与适用性。

原文摘要 · Abstract (English)

This study emphasizes the domain of requirements engineering by applying the SMOTE-Tomek preprocessing technique, combined with stratified K-fold cross-validation, to address class imbalance in the PROMISE dataset. This dataset comprises 969 categorized requirements, classified into functional and non-functional types. The proposed approach enhances the representation of minority classes while maintaining the integrity of validation folds, leading to a notable improvement in classification accuracy. Logistic regression achieved 76.16%, significantly surpassing the baseline of 59.85%. These results highlight the applicability and efficiency of machine learning models as scalable and interpretable solutions.

需求工程分类数据平衡逻辑回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。