arXiv:2506.16586cs.SEcs.AI2025-06被引 10

AI工具可显著提升软件测试效率,但需解决生成结果的可信与可解释性问题。

AI-Driven Tools in Modern Software Quality Assurance: An Assessment of Benefits, Challenges, and Future Directions

  • 用AI生成测试用例并执行端到端回归,实现自动化验证
  • 生成测试用例仅8.3%出现不稳定执行,效果稳定可靠
  • 适合关注AI赋能测试、追求高覆盖率的开发与质量团队

传统质量保障方法难以应对现代分布式软件系统的复杂性、规模及快速迭代需求,且受限于资源,导致质量问题成本高昂。本研究评估了将现代AI工具融入质量保障流程在验证与确认环节的效益、挑战与前景,涵盖探索性测试分析、等价类划分与边界分析、元模型测试、验收标准不一致检测、静态分析、用例生成、单元测试生成、测试套件优化与评估、端到端场景执行等。以企业级应用为样本,通过AI代理执行生成的测试场景进行端到端回归测试,验证了方法可行性。结果显示,生成测试用例仅有8.3%出现不稳定执行,表明其具备显著潜力。然而,实际应用仍面临挑战:大语言模型生成内容语义覆盖不一致、缺乏可解释性,以及倾向于修正测试用例以匹配预期结果,凸显对生成成果和执行结果进行严格验证的必要性。研究表明AI对质量保障具有变革潜力,但需采取战略路径,兼顾局限性并发展配套验证方法。

原文摘要 · Abstract (English)

Traditional quality assurance (QA) methods face significant challenges in addressing the complexity, scale, and rapid iteration cycles of modern software systems and are strained by limited resources available, leading to substantial costs associated with poor quality. The object of this research is the Quality Assurance processes for modern distributed software applications. The subject of the research is the assessment of the benefits, challenges, and prospects of integrating modern AI-oriented tools into quality assurance processes. We performed comprehensive analysis of implications on both verification and validation processes covering exploratory test analyses, equivalence partitioning and boundary analyses, metamorphic testing, finding inconsistencies in acceptance criteria (AC), static analyses, test case generation, unit test generation, test suit optimization and assessment, end to end scenario execution. End to end regression of sample enterprise application utilizing AI-agents over generated test scenarios was implemented as a proof of concept highlighting practical use of the study. The results, with only 8.3% flaky executions of generated test cases, indicate significant potential for the proposed approaches. However, the study also identified substantial challenges for practical adoption concerning generation of semantically identical coverage, "black box" nature and lack of explainability from state-of-the-art Large Language Models (LLMs), the tendency to correct mutated test cases to match expected results, underscoring the necessity for thorough verification of both generated artifacts and test execution results. The research demonstrates AI's transformative potential for QA but highlights the importance of a strategic approach to implementing these technologies, considering the identified limitations and the need for developing appropriate verification methodologies.

AI测试质量保障大模型应用自动化验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。