arXiv:2507.17791cs.LGcs.AI2025-07被引 3

Helix让表格科学数据的机器学习更可复现、可解释,适合非数据科学背景的研究者使用。

Helix 1.0: An Open-Source Framework for Reproducible and Interpretable Machine Learning on Tabular Scientific Data

  • 提供标准化流程,涵盖数据预处理到模型预测全链条
  • 内置语言化解释功能,帮助理解模型决策逻辑
  • 开源免费,支持非专业用户快速上手和协作开发

Helix是一个开源、可扩展的Python框架,旨在为表格型科学数据提供可复现、可解释的机器学习工作流。它解决实验分析过程缺乏透明度的问题,确保数据转换与方法选择等决策全程可记录、可访问、可复现且易懂。系统包含标准化数据预处理、可视化、模型训练、评估、解释、结果检查及未知数据预测等功能模块。为赋能无数据科学训练背景的研究者,Helix提供友好界面,支持设计计算实验、查看结果,并引入一种基于语言术语的新解释方法。项目以MIT许可证发布,可通过GitHub和PyPI获取,支持社区共建,遵循FAIR原则。

原文摘要 · Abstract (English)

Helix is an open-source, extensible, Python-based software framework to facilitate reproducible and interpretable machine learning workflows for tabular data. It addresses the growing need for transparent experimental data analytics provenance, ensuring that the entire analytical process -- including decisions around data transformation and methodological choices -- is documented, accessible, reproducible, and comprehensible to relevant stakeholders. The platform comprises modules for standardised data preprocessing, visualisation, machine learning model training, evaluation, interpretation, results inspection, and model prediction for unseen data. To further empower researchers without formal training in data science to derive meaningful and actionable insights, Helix features a user-friendly interface that enables the design of computational experiments, inspection of outcomes, including a novel interpretation approach to machine learning decisions using linguistic terms all within an integrated environment. Released under the MIT licence, Helix is accessible via GitHub and PyPI, supporting community-driven development and promoting adherence to the FAIR principles.

机器学习可解释性开源工具表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。