arXiv:2504.01356cs.LGcs.SE2025-04

用scikit-learn+MLflow+SHAP打造可追溯的生物医学机器学习快速流程

xML-workFlow: an end-to-end explainable scikit-learn workflow for rapid biomedical experimentation

  • 整合scikit-learn、MLflow与SHAP,构建端到端可复现的建模流程
  • 减少生物医学研究中模型开发与迭代的时间成本,提升可扩展性
  • 适合生物信息学家快速部署机器学习项目,无需重写大量代码

构建和迭代机器学习模型在生物医学研究中常耗时耗力,现有科学代码缺乏可扩展性且难以跨项目复用。xML-workFlow 提供一种快速、稳健且可追踪的端到端工作流,可适配任意机器学习项目,仅需少量代码修改。该流程集成 scikit-learn、MLflow 与 SHAP,显著降低模型开发与迭代所需时间与精力,有效解决生物医学研究中的可扩展性与可复现性挑战。使用本模板可帮助生物信息学家节省开发时间,并使生物医学研究人员更易部署机器学习项目。项目开源地址:https://github.com/MedicalGenomicsLab/xML-workFlow。

原文摘要 · Abstract (English)

Motivation: Building and iterating machine learning models is often a resource-intensive process. In biomedical research, scientific codebases can lack scalability and are not easily transferable to work beyond what they were intended. xML-workFlow addresses this issue by providing a rapid, robust, and traceable end-to-end workflow that can be adapted to any ML project with minimal code rewriting. Results: We show a practical, end-to-end workflow that integrates scikit-learn, MLflow, and SHAP. This template significantly reduces the time and effort required to build and iterate on ML models, addressing the common challenges of scalability and reproducibility in biomedical research. Adapting our template may save bioinformaticians time in development and enables biomedical researchers to deploy ML projects. Availability and implementation: xML-workFlow is available at https://github.com/MedicalGenomicsLab/xML-workFlow.

机器学习生物医学可解释性工作流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。