arXiv:2509.23577cs.DBcs.AI2025-09被引 4

解决机器学习资产难管理、难发现的问题,提升模型与数据利用率。

ML-Asset Management: Curation, Discovery, and Utilization

  • 系统梳理模型、数据集等资产的管理流程,提出全生命周期管理框架。
  • 指出当前资产分散、文档不全、许可混乱等核心问题,影响实际使用效率。
  • 适合研究者与工程师参考,推动真实场景中的资产共享与高效利用。

机器学习(ML)资产如模型、数据集和元数据是现代机器学习工作流的核心。尽管其在实践中呈爆炸式增长,但因文档碎片化、存储孤岛、许可证不一致及缺乏统一发现机制,这些资产往往未被充分利用,使得ML资产管理成为紧迫挑战。本教程全面概述了ML资产在其生命周期中涉及的管理活动,包括整理、发现与利用。我们对ML资产进行分类,梳理主要管理问题,调研前沿技术,并在各阶段识别新兴机遇。同时,强调了可扩展性、溯源关系和统一索引等系统级挑战。通过系统演示,本教程为研究人员和实践者提供可操作的洞察与实用工具,助力在真实世界和特定领域推进ML资产的高效管理。

原文摘要 · Abstract (English)

Machine learning (ML) assets, such as models, datasets, and metadata, are central to modern ML workflows. Despite their explosive growth in practice, these assets are often underutilized due to fragmented documentation, siloed storage, inconsistent licensing, and lack of unified discovery mechanisms, making ML-asset management an urgent challenge. This tutorial offers a comprehensive overview of ML-asset management activities across its lifecycle, including curation, discovery, and utilization. We provide a categorization of ML assets, and major management issues, survey state-of-the-art techniques, and identify emerging opportunities at each stage. We further highlight system-level challenges related to scalability, lineage, and unified indexing. Through live demonstrations of systems, this tutorial equips both researchers and practitioners with actionable insights and practical tools for advancing ML-asset management in real-world and domain-specific settings.

资产管理机器学习数据治理系统设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。