用GPT-4生成软件测试中的等价关系,效果优于GPT-3.5。
Integrating Artificial Intelligence with Human Expertise: An In-depth Analysis of ChatGPT's Capabilities in Generating Metamorphic Relations
- 用GPT-4生成等价关系,结合改进评估框架
- 在9个不同系统上验证,包括含AI/ML的复杂系统
- 人类与AI评估结果对比,展现人机协同潜力
本文深入分析了OpenAI GPT模型在生成和评估等价关系(Metamorphic Relations, MRs)方面的表现,重点考察GPT-4在软件测试环境中的能力。研究首先基于前期标准对GPT-3.5和GPT-4生成的MR进行评估,随后采用改进的评估框架,对GPT-4在9个不同系统(涵盖简单程序到含人工智能/机器学习组件的复杂系统)上生成的MR进行测试。通过自研GPT评估器与人工评估者共同打分,实现自动化与人工评估的直接对比。结果显示,相较于GPT-3.5,GPT-4在生成准确且有用的MR方面显著更优;在新评估标准下,其在各类系统中均展现出高质量生成能力。研究证明GPT-4具备广泛适用的等价关系生成能力,凸显人工智能在软件测试中的潜力,并强调人机协作在该领域的重要价值。
原文摘要 · Abstract (English)
Context: This paper provides an in-depth examination of the generation and evaluation of Metamorphic Relations (MRs) using GPT models developed by OpenAI, with a particular focus on the capabilities of GPT-4 in software testing environments. Objective: The aim is to examine the quality of MRs produced by GPT-3.5 and GPT-4 for a specific System Under Test (SUT) adopted from an earlier study, and to introduce and apply an improved set of evaluation criteria for a diverse range of SUTs. Method: The initial phase evaluates MRs generated by GPT-3.5 and GPT-4 using criteria from a prior study, followed by an application of an enhanced evaluation framework on MRs created by GPT-4 for a diverse range of nine SUTs, varying from simple programs to complex systems incorporating AI/ML components. A custom-built GPT evaluator, alongside human evaluators, assessed the MRs, enabling a direct comparison between automated and human evaluation methods. Results: The study finds that GPT-4 outperforms GPT-3.5 in generating accurate and useful MRs. With the advanced evaluation criteria, GPT-4 demonstrates a significant ability to produce high-quality MRs across a wide range of SUTs, including complex systems incorporating AI/ML components. Conclusions: GPT-4 exhibits advanced capabilities in generating MRs suitable for various applications. The research underscores the growing potential of AI in software testing, particularly in the generation and evaluation of MRs, and points towards the complementarity of human and AI skills in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。