用大模型生成需求工程数据,提升缺陷识别效果。
Synthline: A Product Line Approach for Synthetic Requirements Engineering Data Generation using Large Language Models
- 基于产品线思想,用大模型系统生成合成需求数据。
- 混合真实与合成数据,召回率提升2倍,精确率最高增85%。
- 适合研究需求工程数据稀缺问题的学者和开发者。
现代需求工程(RE)依赖自然语言处理与机器学习技术,但其效果受限于高质量数据集的匮乏。本文提出Synthline,一种基于产品线(PL)的方法,利用大语言模型系统生成用于分类任务的合成RE数据。通过在缺陷识别场景下的实证评估,我们分析了生成数据的多样性及其对下游模型训练的效用。结果表明,尽管合成数据多样性低于真实数据,但仍可作为有效的训练资源。更重要的是,将合成数据与真实数据结合使用,能显著提升模型性能:相比仅用真实数据训练的模型,混合方法在精确率上最高提升85%,召回率提升2倍。这些发现证明了基于产品线的合成数据生成在缓解需求工程数据稀缺问题上的潜力。我们公开了实现代码与生成数据集,以支持研究复现与领域发展。
原文摘要 · Abstract (English)
While modern Requirements Engineering (RE) heavily relies on natural language processing and Machine Learning (ML) techniques, their effectiveness is limited by the scarcity of high-quality datasets. This paper introduces Synthline, a Product Line (PL) approach that leverages Large Language Models to systematically generate synthetic RE data for classification-based use cases. Through an empirical evaluation conducted in the context of using ML for the identification of requirements specification defects, we investigated both the diversity of the generated data and its utility for training downstream models. Our analysis reveals that while synthetic datasets exhibit less diversity than real data, they are good enough to serve as viable training resources. Moreover, our evaluation shows that combining synthetic and real data leads to substantial performance improvements. Specifically, hybrid approaches achieve up to 85% improvement in precision and a 2x increase in recall compared to models trained exclusively on real data. These findings demonstrate the potential of PL-based synthetic data generation to address data scarcity in RE. We make both our implementation and generated datasets publicly available to support reproducibility and advancement in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。