arXiv:2506.16051cs.LGcs.DB2025-06被引 2

构建可复现的协作机器学习数据基础设施

From Data to Decision: Data-Centric Infrastructure for Reproducible ML in Collaborative eScience

  • 以数据为中心,定义六类结构化产物
  • 实现模型迭代全过程的版本化与可追溯
  • 适合需要长期协作与结果复现的科研团队

机器学习的可复现性仍是协作型eScience中的核心挑战,尤其在数据、特征和模型持续迭代的项目中。当前工作流动态但碎片化,依赖非正式的数据共享、临时脚本和松散连接的工具,阻碍了透明性、可复现性和实验的长期适应性。本文提出一种面向生命周期的以数据为中心的复现框架,包含六个结构化产物:数据集(Dataset)、特征(Feature)、工作流(Workflow)、执行记录(Execution)、资产(Asset)和受控词汇表(Controlled Vocabulary)。这些产物形式化地表达数据、代码与决策之间的关系,使机器学习实验能够被版本化、可解释且全程可追溯。该方法在青光眼检测的临床机器学习案例中得到验证,展示了系统如何支持迭代探索,提升可复现性,并保留协作决策的来源信息。

原文摘要 · Abstract (English)

Reproducibility remains a central challenge in machine learning (ML), especially in collaborative eScience projects where teams iterate over data, features, and models. Current ML workflows are often dynamic yet fragmented, relying on informal data sharing, ad hoc scripts, and loosely connected tools. This fragmentation impedes transparency, reproducibility, and the adaptability of experiments over time. This paper introduces a data-centric framework for lifecycle-aware reproducibility, centered around six structured artifacts: Dataset, Feature, Workflow, Execution, Asset, and Controlled Vocabulary. These artifacts formalize the relationships between data, code, and decisions, enabling ML experiments to be versioned, interpretable, and traceable over time. The approach is demonstrated through a clinical ML use case of glaucoma detection, illustrating how the system supports iterative exploration, improves reproducibility, and preserves the provenance of collaborative decisions across the ML lifecycle.

可复现性协作科研数据管理机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。