arXiv:2508.17900cs.SEcs.AI2025-08被引 1

为AI软件缺陷分析设计新框架,识别学习与思考阶段的高风险问题。

A Defect Classification Framework for AI-Based Software Systems (AI-ODC)

  • 基于ODC框架扩展数据、学习、思考三维度分类体系。
  • 学习阶段缺陷最常见,且与高严重性显著相关。
  • 适合关注AI系统质量保障与缺陷预防的研究者与工程师。

人工智能在日常生活和制造业中广泛应用,其系统多以软件形式实现,缺陷分析是保障质量的关键环节。然而现有缺陷分析模型难以捕捉AI系统的独特属性。本文提出AI-ODC框架,基于正交缺陷分类(ODC)范式,引入数据、学习、思考三类新维度,调整属性、新增严重性等级,并将影响区域替换为与AI特性相关的特征。通过在公开的机器学习缺陷数据集上应用该框架,一维与二维分析结果显示:学习阶段缺陷最为普遍,且与高严重性显著相关;而思考阶段缺陷对可信度与准确性影响尤为突出。该结果表明AI-ODC能有效识别高风险缺陷类别,支持针对性的质量保障措施。

原文摘要 · Abstract (English)

Artificial Intelligence has gained a lot of attention recently, it has been utilized in several fields ranging from daily life activities, such as responding to emails and scheduling appointments, to manufacturing and automating work activities. Artificial Intelligence systems are mainly implemented as software solutions, and it is essential to discover and remove software defects to assure its quality using defect analysis which is one of the major activities that contribute to software quality. Despite the proliferation of AI-based systems, current defect analysis models fail to capture their unique attributes. This paper proposes a framework inspired by the Orthogonal Defect Classification (ODC) paradigm and enables defect analysis of Artificial Intelligence systems while recognizing its special attributes and characteristics. This study demonstrated the feasibility of modifying ODC for AI systems to classify its defects. The ODC was adjusted to accommodate the Data, Learning, and Thinking aspects of AI systems which are newly introduced classification dimensions. This adjustment involved the introduction of an additional attribute to the ODC attributes, the incorporation of a new severity level, and the substitution of impact areas with characteristics pertinent to AI systems. The framework was showcased by applying it to a publicly available Machine Learning bug dataset, with results analyzed through one-way and two-way analysis. The case study indicated that defects occurring during the Learning phase were the most prevalent and were significantly linked to high-severity classifications. In contrast, defects identified in the Thinking phase had a disproportionate effect on trustworthiness and accuracy. These findings illustrate AIODC's capability to identify high-risk defect categories and inform focused quality assurance measures.

缺陷分析AI质量机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。