用合成数据提升大模型生成需求文档的评估精度
ReqGenX: An Empirical Study of Atomic Decomposition, Artifact Regeneration, and Reconstruction for Legacy SRS Documents

- 将旧需求文档拆解为原子语句,再重构为可追溯的中间产物
- 生成结果在语义对齐度上达0.96以上,97%通过人工评测
- 适合做需求生成模型评估的研究者和工程团队
背景:自动化软件需求规格(SRS)生成的评估困难,因缺乏源需求、中间采集产物与生成规格之间的细粒度可追溯性。目标:研究是否可将遗留SRS文档转化为支持细粒度评估的可追溯合成预SRS产物。方法:采用ReqGenX控制流程,将SRS章节分解为基于源的原子陈述,通过多大模型投票分配至标准类产物类型,并使用约束提示与迭代判别引导优化生成。在七个PURE SRS文档上评估了接地性、质量、信息保留及下游重建效果。结果:生成原子语义忠实,中位对齐得分0.96–0.99,Prometheus评分4.34–4.85;产物强关联源原子,对齐得分0.80–0.94,人工评测通过率近100%;严格评估下通过率54.8%–97.1%。下游重建案例中,基于产物的原子在生成SRS中仍可恢复,SBERT均值0.69–0.75,对齐得分中位0.76–0.84。结论:可追溯的合成预SRS产物能支撑更精细的LLM-SRS生成评估,同时揭示忠实性、信息保留与产物完整性间的权衡。
原文摘要 · Abstract (English)
Background: Evaluating automated Software Requirements Specification (SRS) generation is challenging because few datasets provide fine-grained traceability between source requirements, intermediate elicitation artifacts, and generated specifications. Aims: We aim to study whether legacy SRS documents can be transformed into traceable synthetic pre-SRS artifacts that support fine-grained evaluation of LLM-based SRS generation. Method: We conduct an empirical study using ReqGenX, a controlled pipeline that decomposes SRS sections into source-grounded atomic statements, routes atoms to standards-inspired artifact types through multi-LLM plurality voting, and generates artifacts using constrained prompts with iterative judge-guided refinement. We evaluate ReqGenX on seven PURE SRS documents using grounding, quality, information retention, and downstream reconstruction analyses. Results: ReqGenX produces faithful and usable atoms, with median AlignScore values typically between 0.96 and 0.99 and Prometheus scores ranging from 4.34 to 4.85. Generated artifacts remain strongly grounded in their source atoms, with AlignScore values typically between 0.80--0.94 and judge pass rates near 100%; stricter Prometheus evaluation yields pass rates from 54.8% to 97.1%. In a downstream SRS reconstruction case study, artifact-backed atoms remain recoverable from generated SRSs, with SBERT means between 0.69 and 0.75 and AlignScore medians between 0.76 and 0.84. Conclusions: Traceable synthetic pre-SRS artifacts can support more fine-grained evaluation of LLM-based SRS generation, while exposing tradeoffs among faithfulness, information retention, and artifact completeness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。