arXiv:2604.10965stat.COcs.LG2026-04被引 1

解决生物医学机器学习中的数据泄露问题,确保模型评估更真实可靠。

bioLeak: Leakage-Aware Modeling and Diagnostics for Machine Learning in R

  • 设计泄漏感知的交叉验证流程与预处理策略
  • 模拟显示泄露会显著夸大模型性能表现
  • 适合生物医学数据建模者用于排查分析偏差

数据泄露是生物医学机器学习研究中导致乐观偏差的常见根源。针对重复测量、研究间异质性、批次效应或时间依赖性的数据,传统的行级交叉验证和全局预处理方法往往不适用。本文介绍 bioLeak,一个用于构建泄漏感知重采样工作流并审计拟合模型中常见泄漏机制的 R 包。该包支持泄漏感知的分组构造、仅训练集预处理、交叉验证建模、嵌套超参数调优、事后泄漏审计及 HTML 报告生成。支持二分类、多分类、回归与生存分析任务,配备任务特定指标和 S4 容器管理分组、拟合、审计与膨胀汇总结果。仿真结果显示,在受控泄漏机制下模型性能明显虚高;案例研究表明,有防护与有泄露的管道在多研究转录组数据上得出截然不同的结论。全文强调软件设计、可复现工作流及诊断输出的解释。

原文摘要 · Abstract (English)

Data leakage remains a recurrent source of optimistic bias in biomedical machine learning studies. Standard row-wise cross-validation and globally estimated preprocessing steps are often inappropriate for data with repeated measurements, study-level heterogeneity, batch effects, or temporal dependencies. This paper describes bioLeak, an R package for constructing leakage-aware resampling workflows and for auditing fitted models for common leakage mechanisms. The package provides leakage-aware split construction, train-fold-only preprocessing, cross-validated model fitting, nested hyperparameter tuning, post hoc leakage audits, and HTML reporting. The implementation supports binary classification, multiclass classification, regression, and survival analysis, with task-specific metrics and S4 containers for splits, fits, audits, and inflation summaries. The simulation artifacts show how apparent performance changes under controlled leakage mechanisms, and the case study illustrates how guarded and leaky pipelines can yield materially different conclusions on multi-study transcriptomic data. The emphasis throughout is on software design, reproducible workflows, and interpretation of diagnostic output.

机器学习数据泄露生物信息R语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。