用可量化指标自动评估生成的测试文档质量,提升可靠性。
Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
- 用反向生成+评分规则迭代优化,提升文档清晰度和完整性。
- 12个项目的实验显示,劣质输入经处理后可达到高质量输出。
- 适合想在自动化中保障质量的工程团队使用。
大型语言模型(LLMs)正在改变质量工程(QE),实现需求、测试用例及行为驱动开发(BDD)场景等文档的自动化生成。然而,确保输出质量仍是挑战。本文提出一种系统性技术,通过可量化的指标对QE文档进行基准测试与评估。该方法结合LLM生成、反向生成以及基于清晰性、完整性、一致性、可测试性评分规则的迭代优化。在12个实际项目上的实验表明,反向生成的文档可超越低质量输入,并在输入质量高时维持高水平输出。该框架实现了可扩展、可靠的QE文档验证,弥合了自动化与责任之间的鸿沟。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are transforming Quality Engineering (QE) by automating the generation of artefacts such as requirements, test cases, and Behavior Driven Development (BDD) scenarios. However, ensuring the quality of these outputs remains a challenge. This paper presents a systematic technique to baseline and evaluate QE artefacts using quantifiable metrics. The approach combines LLM-driven generation, reverse generation , and iterative refinement guided by rubrics technique for clarity, completeness, consistency, and testability. Experimental results across 12 projects show that reverse-generated artefacts can outperform low-quality inputs and maintain high standards when inputs are strong. The framework enables scalable, reliable QE artefact validation, bridging automation with accountability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。