系统梳理MLOps中保障模型可信部署的实践与挑战。
Towards Trustworthy Machine Learning in Production: An Overview of the Robustness in MLOps Approach
- 提出构建鲁棒MLOps的实用技术路径
- 总结现有研究对生产环境鲁棒性的支持程度
- 适合关注AI落地可靠性的研发与工程人员
人工智能,尤其是机器学习(ML),正深刻影响日常生活。近年来,研究人员和从业者提出了确保系统做出可靠、可信决策的原则与指南。传统ML系统依赖历史数据提取特征并训练模型完成任务,但在实际部署中,持续运行于真实环境面临根本性挑战。为此,机器学习运维(MLOps)成为标准化部署方案的潜在解法。尽管MLOps在优化流程方面取得显著成效,但如何全面定义鲁棒的MLOps体系仍受广泛关注。本文综述了MLOps系统的可信性属性,重点强调实现鲁棒MLOps的技术实践;系统调研现有研究在生产环境中提升模型鲁棒性的方法;回顾可用工具与软件对鲁棒性支持情况;最后提出当前开放挑战及未来方向。旨在为从事实际AI应用的研究者与从业者提供生产环境部署鲁棒解决方案的全景视图。
原文摘要 · Abstract (English)
Artificial intelligence (AI), and especially its sub-field of Machine Learning (ML), are impacting the daily lives of everyone with their ubiquitous applications. In recent years, AI researchers and practitioners have introduced principles and guidelines to build systems that make reliable and trustworthy decisions. From a practical perspective, conventional ML systems process historical data to extract the features that are consequently used to train ML models that perform the desired task. However, in practice, a fundamental challenge arises when the system needs to be operationalized and deployed to evolve and operate in real-life environments continuously. To address this challenge, Machine Learning Operations (MLOps) have emerged as a potential recipe for standardizing ML solutions in deployment. Although MLOps demonstrated great success in streamlining ML processes, thoroughly defining the specifications of robust MLOps approaches remains of great interest to researchers and practitioners. In this paper, we provide a comprehensive overview of the trustworthiness property of MLOps systems. Specifically, we highlight technical practices to achieve robust MLOps systems. In addition, we survey the existing research approaches that address the robustness aspects of ML systems in production. We also review the tools and software available to build MLOps systems and summarize their support to handle the robustness aspects. Finally, we present the open challenges and propose possible future directions and opportunities within this emerging field. The aim of this paper is to provide researchers and practitioners working on practical AI applications with a comprehensive view to adopt robust ML solutions in production environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。