让AI生成的测试用例更可靠,通过治理框架降低风险。
Governance Controls for AI-Generated Test Artifacts in Autonomous Software Testing

- 引入治理感知框架,在测试流程中加入合规与风险评估
- 风险降低89.6%,测试可靠性达96.5%,合规准确率94.2%
- 适合关注AI测试安全与可解释性的研发团队使用
人工智能和大语言模型在自主软件测试中应用日益广泛,但生成的测试用例常存在幻觉、合规违规、安全风险及可解释性差等问题。为提升其可靠性、透明度与可信度,本文提出治理感知自主测试框架(GATF),在自主测试生命周期中引入治理验证、可解释性分析、概率风险评估、合规监控与审计治理。在Defects4J和PROMISE数据集上的实验表明,该框架将治理相关风险降低89.6%,实现94.3%的治理准确性、96.5%的测试用例可靠性、94.2%的合规准确率和90.8%的可解释性表现。结果表明,具备治理意识的自主测试系统显著优于传统AI测试方法,架构具备可扩展性与可靠性,可为软件测试提供安全环境。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) and Large Language Models (LLMs) are increasingly used in autonomous software testing; however, AI-generated test artifacts often suffer from hallucinations, compliance violations, security risks, and limited explainability. To enhance the reliability, transparency, and trustworthiness of AI-generated testing artifacts, this research introduces the concept of Governance-Aware Autonomous Testing Framework (GATF). The framework extends the autonomous testing lifecycle with governance validation, explainability analysis, probabilistic risk assessment, compliance monitoring, as well as audit governance. Experiments were performed with Defects4J and PROMISE software engineering datasets. The proposed framework successfully reduced the governance-related risks by 89.6% and demonstrated 94.3% accuracy in governance, 96.5% artifact reliability, 94.2% compliance accuracy, and 90.8% explainability performance. The results show that autonomous testing systems that are governance-aware can significantly enhance the reliability, transparency, and operational security of autonomous testing systems in comparison to conventional AI-based testing systems. The proposed architecture is scalable and reliable and provides a safe environment for software testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。