数据预处理常掩盖医学模型真相,反而降低可解释性。
Common Steps in Machine Learning Might Hinder The Explainability Aims in Medicine
- 分析常见预处理步骤如何影响医学模型可解释性
- 不当处理缺失值和异常值会遮蔽新发现
- 适合关注模型公平性与临床意义的医疗AI研究者
数据预处理是机器学习中提升模型性能和减少运行时间的重要步骤,包括处理缺失值、检测并移除异常值、数据增强、降维、归一化及控制混杂变量影响等。尽管这些步骤能提高模型准确率,但在医学领域若未审慎设计,可能损害模型的可解释性。不当的缺失值或异常值处理会掩盖潜在医学发现;某些预处理可能导致模型对不同群体决策不公平;此外,特征被转化为无单位且临床无意义的数值,使结果难以解释。本文系统讨论了常见的数据预处理步骤及其对模型可解释性和可理解性的影响,并提出在保持性能的同时提升可解释性的可行方案。
原文摘要 · Abstract (English)
Data pre-processing is a significant step in machine learning to improve the performance of the model and decreases the running time. This might include dealing with missing values, outliers detection and removing, data augmentation, dimensionality reduction, data normalization and handling the impact of confounding variables. Although it is found the steps improve the accuracy of the model, but they might hinder the explainability of the model if they are not carefully considered especially in medicine. They might block new findings when missing values and outliers removal are implemented inappropriately. In addition, they might make the model unfair against all the groups in the model when making the decision. Moreover, they turn the features into unitless and clinically meaningless and consequently not explainable. This paper discusses the common steps of the data preprocessing in machine learning and their impacts on the explainability and interpretability of the model. Finally, the paper discusses some possible solutions that improve the performance of the model while not decreasing its explainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。