对比BERT与随机森林在跨项目安全漏洞报告预测中的表现
Security Bug Report Prediction Within and Across Projects: A Comparative Study of BERT and Random Forest
- 用随机森林和BERT模型对比预测安全漏洞报告
- 跨项目预测中BERT达到62%的G度量,显著优于随机森林
- 结合非安全漏洞数据可提升BERT性能,但会降低随机森林表现
早期发现安全漏洞报告(SBR)对预防系统漏洞、保障可靠性至关重要。尽管已有机器学习模型用于预测SBR,其性能仍有提升空间。本研究全面比较了BERT与随机森林(RF)在预测SBR上的表现。结果显示,在项目内预测中,RF平均G度量比BERT高出34%。仅引入其他项目的SBR可提升两模型的平均性能。然而,加入安全与非安全漏洞报告后,RF平均性能降至46%,而BERT性能提升至最佳值66%,超过RF。在跨项目预测中,BERT取得62%的显著高G度量,远超随机森林。
原文摘要 · Abstract (English)
Early detection of security bug reports (SBRs) is crucial for preventing vulnerabilities and ensuring system reliability. While machine learning models have been developed for SBR prediction, their predictive performance still has room for improvement. In this study, we conduct a comprehensive comparison between BERT and Random Forest (RF), a competitive baseline for predicting SBRs. The results show that RF outperforms BERT with a 34% higher average G-measure for within-project predictions. Adding only SBRs from various projects improves both models' average performance. However, including both security and nonsecurity bug reports significantly reduces RF's average performance to 46%, while boosts BERT to its best average performance of 66%, surpassing RF. In cross-project SBR prediction, BERT achieves a remarkable 62% G-measure, which is substantially higher than RF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。