构建支持多任务的机器学习数据虚拟化系统,提升研发效率。
Data Virtualization for Machine Learning
- 通过数据虚拟化技术统一管理多任务机器学习流程中的中间数据
- 支撑六个应用、多个工作流,可扩展至未来更多任务
- 适合需要高效协同和长期迭代的ML团队
当前,机器学习团队同时运行多个面向不同应用的实验流程,每个流程涉及大量实验、迭代与协作,从数据准备到模型部署通常耗时数月甚至数年。组织层面需存储、处理并维护大量中间数据。数据虚拟化成为支撑机器学习工作流的核心基础设施技术。本文介绍了一套数据虚拟化服务的设计与实现,重点阐述其服务架构与运维机制。该基础设施目前已支持六个机器学习应用,每个应用包含多个机器学习工作流,未来可继续扩展以支持更多应用与工作流。
原文摘要 · Abstract (English)
Nowadays, machine learning (ML) teams have multiple concurrent ML workflows for different applications. Each workflow typically involves many experiments, iterations, and collaborative activities and commonly takes months and sometimes years from initial data wrangling to model deployment. Organizationally, there is a large amount of intermediate data to be stored, processed, and maintained. \emph{Data virtualization} becomes a critical technology in an infrastructure to serve ML workflows. In this paper, we present the design and implementation of a data virtualization service, focusing on its service architecture and service operations. The infrastructure currently supports six ML applications, each with more than one ML workflow. The data virtualization service allows the number of applications and workflows to grow in the coming years.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。