arXiv:2511.21354cs.LG2025-11被引 1

为科学领域机器学习实验提供可复现的规范流程

Best Practices for Machine Learning Experimentation in Scientific Applications

  • 构建从数据准备到评估的全流程标准化方法
  • 提出LOR和COS指标以量化过拟合与结果不稳定性
  • 适合希望提升实验可信度的科研人员参考

机器学习在科学研究中的应用日益广泛,但实验设计与记录质量直接影响结果的可靠性。不当的基线设置、不一致的数据预处理或不足的验证可能导致模型性能结论失真。本文提出一套面向科学应用的机器学习实验实践指南,聚焦可复现性、公平比较与透明报告。我们梳理了从数据准备到模型选择与评估的分步工作流,并引入对过拟合及验证折间不稳定性敏感的评估指标,包括对数过拟合比(Logarithmic Overfitting Ratio, LOR)和综合过拟合得分(Composite Overfitting Score, COS)。通过推荐做法与示例报告格式,旨在帮助研究人员建立稳健基线,基于可靠证据从机器学习模型中得出有效科学洞见。

原文摘要 · Abstract (English)

Machine learning (ML) is increasingly adopted in scientific research, yet the quality and reliability of results often depend on how experiments are designed and documented. Poor baselines, inconsistent preprocessing, or insufficient validation can lead to misleading conclusions about model performance. This paper presents a practical and structured guide for conducting ML experiments in scientific applications, focussing on reproducibility, fair comparison, and transparent reporting. We outline a step-by-step workflow, from dataset preparation to model selection and evaluation, and propose metrics that account for overfitting and instability across validation folds, including the Logarithmic Overfitting Ratio (LOR) and the Composite Overfitting Score (COS). Through recommended practices and example reporting formats, this work aims to support researchers in establishing robust baselines and drawing valid evidence-based insights from ML models applied to scientific problems.

机器学习实验规范可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。