梳理40+条病历数据提取难题,给出临床预测模型实用指南
Challenges and recommendations for Electronic Health Records data extraction and preparation for dynamic prediction modelling in hospitalized patients -- a practical guide
- 按患者定义、结局判定、特征工程、数据清洗四类归纳挑战
- 提出可操作建议,提升动态预测模型的数据质量与可信度
- 适合医疗数据工程师与临床研究者参考,降低建模偏差
利用电子健康记录(EHR)数据进行动态预测建模近年来受到广泛关注。此类模型的可靠性与可信度高度依赖于底层数据质量,而数据质量又在很大程度上由模型开发前的数据提取与准备阶段决定。本文识别出该阶段逾40项常见挑战,并针对这些挑战提出具体可行的改进建议。挑战被归纳为四大类别:队列定义、结局定义、特征工程与数据清洗。本综述为数据提取工程师与研究人员提供了一套实用指导,旨在推广最佳实践,提升动态预测模型在真实临床环境中的质量与应用价值。
原文摘要 · Abstract (English)
Dynamic predictive modelling using electronic health record (EHR) data has gained significant attention in recent years. The reliability and trustworthiness of such models depend heavily on the quality of the underlying data, which is, in part, determined by the stages preceding the model development: data extraction from EHR systems and data preparation. In this article, we identified over forty challenges encountered during these stages and provide actionable recommendations for addressing them. These challenges are organized into four categories: cohort definition, outcome definition, feature engineering, and data cleaning. This comprehensive list serves as a practical guide for data extraction engineers and researchers, promoting best practices and improving the quality and real-world applicability of dynamic prediction models in clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。