用AI从自然语言需求生成测试用例,解决传统方法耗时难题。
AI-Driven Test Case Generation from Natural Language Requirements: A Survey of Techniques and Research Gaps
- 基于AI与NLP技术,将自然语言需求自动转化为测试用例。
- 现有方法无法同时满足自动化、抗模糊、可追溯等六项质量标准。
- 适合关注AI辅助测试、软件工程自动化方向的研究者。
软件测试是验证系统是否满足需求的关键环节,但仍是开发中最耗时且成本最高的活动之一。基于需求的测试用例生成可在早期从需求文档中提取测试用例,但直接从自然语言生成面临固有的歧义和不精确问题。近年来,人工智能、自然语言处理(NLP)及大语言模型(LLMs)的进步使该流程自动化成为可能,但也引入了幻觉、可追溯性降低和评估不一致等新风险。本综述围绕四个研究问题展开:有哪些AI与NLP技术用于从自然语言需求生成测试用例;有哪些工具与框架支持这些方法;如何评估生成的测试用例;现存哪些研究空白。依据Kitchenham与Charters的系统综述指南,检索2000–2025年主要学术数据库,经严格筛选后确定21篇核心研究。文献按三个演进阶段组织,揭示当前方法均无法同时满足六项关键质量维度:自动化、歧义处理、领域适用性、可追溯性、评估充分性与幻觉控制。本文贡献包括:三阶段演进综述、六维度差距分析,以及四项针对幻觉、可追溯性、复杂度敏感性与合规性的可操作研究建议。
原文摘要 · Abstract (English)
Software testing is critical for verifying that systems meet specified requirements, yet remains among the most time-consuming and expensive activities in development. Requirements-based test generation allows test cases to be derived early from requirements artifacts, but generating them directly from natural language is challenging due to inherent ambiguity and imprecision. Recent advances in AI, natural language processing (NLP), and large language models (LLMs) have made automating this pipeline increasingly feasible, while introducing new risks including hallucination, reduced traceability, and inconsistent evaluation. This survey addresses four research questions: what AI and NLP techniques have been proposed for generating test cases from natural language requirements; what tools and frameworks support these approaches; how generated test cases are evaluated; and what research gaps remain. Following Kitchenham and Charters' systematic review guidelines, we searched major scholarly databases spanning 2000-2025 and, after applying strict inclusion criteria, identified 21 primary studies. The literature is organized into three evolutionary eras, revealing that no existing approach simultaneously satisfies six key quality dimensions: automation, ambiguity handling, domain applicability, traceability, evaluation thoroughness, and hallucination control. The survey makes three main contributions: a three-era evolutionary synthesis of AI-based test generation; a six-criteria gap analysis showing no current approach fully addresses all quality dimensions; and four actionable research guidelines targeting hallucination, traceability, complexity sensitivity, and compliance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。