自动处理医疗数据表征,一键生成高效基线模型。
MEDS-Tab: Automated tabularization and baseline methods for MEDS datasets
- 基于MEDS框架自动处理时序医疗数据,支持大规模特征提取。
- 可在数十万特征、数亿临床事件上快速生成高质量XGBoost基线。
- 适合希望快速验证算法的医疗AI研究者,降低人工成本。
为实现结构化电子健康记录(EHR)数据上机器学习解决方案的有效、可靠与可扩展开发,需要能高效、高性能地生成多样监督学习任务的高质量基线模型。过去,此类基线模型的构建高度依赖人工:研究人员需手动决定对原始纵向数据进行特征工程与表格式转换,并训练监督模型以获得可比基线结果,仅针对单一任务和数据集。本工作依托MEDS框架在核心数据标准化方面的进展,显著简化并加速了不规则采样时间序列数据的表征过程,使研究者能够自动、可扩展地对纵向EHR数据进行特征化与表格式转换,覆盖数十万项特征、数亿临床事件,以及多种窗口范围与聚合策略。随后,系统自动利用这些表格数据以高度计算效率的方式生成高水准的XGBoost基线模型。该系统可处理远超现有工具规模的数据集,使任何具备MEDS格式数据的研究者均可立即开展各类任务的可靠、高性能基线预测,几乎无需人工干预。这将极大提升医疗问题中复杂机器学习方案的可靠性、可复现性与开发便捷性。
原文摘要 · Abstract (English)
Effective, reliable, and scalable development of machine learning (ML) solutions for structured electronic health record (EHR) data requires the ability to reliably generate high-quality baseline models for diverse supervised learning tasks in an efficient and performant manner. Historically, producing such baseline models has been a largely manual effort--individual researchers would need to decide on the particular featurization and tabularization processes to apply to their individual raw, longitudinal data; and then train a supervised model over those data to produce a baseline result to compare novel methods against, all for just one task and one dataset. In this work, powered by complementary advances in core data standardization through the MEDS framework, we dramatically simplify and accelerate this process of tabularizing irregularly sampled time-series data, providing researchers the ability to automatically and scalably featurize and tabularize their longitudinal EHR data across tens of thousands of individual features, hundreds of millions of clinical events, and diverse windowing horizons and aggregation strategies, all before ultimately leveraging these tabular data to automatically produce high-caliber XGBoost baselines in a highly computationally efficient manner. This system scales to dramatically larger datasets than tabularization tools currently available to the community and enables researchers with any MEDS format dataset to immediately begin producing reliable and performant baseline prediction results on various tasks, with minimal human effort required. This system will greatly enhance the reliability, reproducibility, and ease of development of powerful ML solutions for health problems across diverse datasets and clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。