arXiv:2602.05043cs.SEcs.AI2026-02中稿 · CAIN 2026, the 5th…

为机器学习组件建立可落地的质量评估框架,解决模型与系统需求脱节问题。

Quality Model for Machine Learning Components

  • 区分组件级与系统级质量属性,聚焦模型开发者的可操作需求
  • 通过调研验证模型有效性,被集成至开源测试工具并实际应用
  • 适合模型开发者与系统团队协作时明确质量要求和测试重点

尽管机器学习(ML)应用日益广泛,但研究显示许多原型未能进入生产环境,且测试仍主要关注模型性能等指标,忽视了系统层面的需求,如吞吐量、资源消耗或鲁棒性。这种片面的测试方式导致模型集成、部署和运维失败。传统软件质量标准如ISO 25010提供结构化框架,而新标准ISO 25059虽针对AI系统,却将系统属性与模型属性混在一起,对模型开发者无实际帮助。本文提出一个专为机器学习组件设计的质量模型,支持需求获取与协商,提供开发者与系统利益相关者共同的语言,明确系统衍生需求并指导测试重点。该模型经调查验证,参与者普遍认可其相关性与价值,并已成功集成至开源机器学习组件测试与评估工具中,展现实际应用潜力。

原文摘要 · Abstract (English)

Despite increased adoption and advances in machine learning (ML), there are studies showing that many ML prototypes do not reach the production stage and that testing is still largely limited to testing model properties, such as model performance, without considering requirements derived from the system it will be a part of, such as throughput, resource consumption, or robustness. This limited view of testing leads to failures in model integration, deployment, and operations. In traditional software development, quality models such as ISO 25010 provide a widely used structured framework to assess software quality, define quality requirements, and provide a common language for communication with stakeholders. A newer standard, ISO 25059, defines a more specific quality model for AI systems. However, a problem with this standard is that it combines system attributes with ML component attributes, which is not helpful for a model developer, as many system attributes cannot be assessed at the component level. In this paper, we present a quality model for ML components that serves as a guide for requirements elicitation and negotiation and provides a common vocabulary for ML component developers and system stakeholders to agree on and define system-derived requirements and focus their testing efforts accordingly. The quality model was validated through a survey in which the participants agreed with its relevance and value. The quality model has been successfully integrated into an open-source tool for ML component testing and evaluation demonstrating its practical application.

质量评估ML组件测试框架系统集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。