针对小样本脑成像数据,提出可复现的抗偏差机器学习框架。
A Reproducible Framework for Bias-Resistant Machine Learning on Small-Sample Neuroimaging Data
- 融合领域知识特征工程与嵌套交叉验证,避免结果过拟合
- 在嵌套交叉验证下实现0.660±0.068的平衡准确率
- 适合小样本生物医学建模,强调可解释性与结果可信度
我们提出一种可复现、抗偏差的机器学习框架,整合领域知识特征工程、嵌套交叉验证和校准决策阈值优化,适用于小样本神经影像数据。传统交叉验证因重复使用相同划分进行模型选择与性能评估,导致结果过于乐观,影响可复现性与泛化能力。在高维结构性MRI数据集(深部脑刺激认知结局)上,该框架通过重要性导向排序选出紧凑可解释子集,实现嵌套交叉验证下0.660±0.068的平衡准确率。结合可解释性与无偏评估,本研究为数据受限的生物医学领域提供了可靠的机器学习通用范式。
原文摘要 · Abstract (English)
We introduce a reproducible, bias-resistant machine learning framework that integrates domain-informed feature engineering, nested cross-validation, and calibrated decision-threshold optimization for small-sample neuroimaging data. Conventional cross-validation frameworks that reuse the same folds for both model selection and performance estimation yield optimistically biased results, limiting reproducibility and generalization. Demonstrated on a high-dimensional structural MRI dataset of deep brain stimulation cognitive outcomes, the framework achieved a nested-CV balanced accuracy of 0.660\,$\pm$\,0.068 using a compact, interpretable subset selected via importance-guided ranking. By combining interpretability and unbiased evaluation, this work provides a generalizable computational blueprint for reliable machine learning in data-limited biomedical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。