CRISP统一处理多中心重症数据,让研究者省下数月预处理时间。
The CRITICAL Records Integrated Standardization Pipeline (CRISP): End-to-End Processing of Large-scale Multi-institutional OMOP CDM Data
- 构建端到端数据标准化流水线,支持跨机构医疗数据整合
- 19.5亿条记录在一天内完成处理,支持标准硬件运行
- 提供基准模型与文档,助力临床AI研究快速启动
现有重症监护电子病历数据集如MIMIC和eICU虽推动了临床AI发展,但新发布的CRITICAL数据集实现了规模与多样性的突破——涵盖来自4个地理分布不同的CTSA机构的371,365名患者,共计19.5亿条记录。该数据集独特优势在于完整覆盖患者从院前、入院、重症监护到出院后的全周期诊疗过程,包含住院与门诊场景。这种多中心、纵向视角为构建可泛化预测模型和推进健康公平研究提供了变革性机遇。然而,多机构数据异构性强,术语体系差异大,导致数据整合复杂度高。为此我们提出CRISP,通过四步流程释放该资源潜力:(1)透明的数据质量管理与完整审计追踪;(2)将异构医学术语映射至统一的SNOMED-CT标准,实现去重与单位标准化;(3)模块化架构支持并行优化,在普通计算设备上完成全流程处理耗时小于1天;(4)提供覆盖多种临床预测任务的基准模型,建立可复现的性能评估标准。通过提供处理流程、基准实现与详尽转换文档,CRISP使研究人员节省数月预处理时间,降低使用门槛,真正实现大规模多中心重症数据的普惠可用。
原文摘要 · Abstract (English)
While existing critical care EHR datasets such as MIMIC and eICU have enabled significant advances in clinical AI research, the CRITICAL dataset opens new frontiers by providing extensive scale and diversity -- containing 1.95 billion records from 371,365 patients across four geographically diverse CTSA institutions. CRITICAL's unique strength lies in capturing full-spectrum patient journeys, including pre-ICU, ICU, and post-ICU encounters across both inpatient and outpatient settings. This multi-institutional, longitudinal perspective creates transformative opportunities for developing generalizable predictive models and advancing health equity research. However, the richness of this multi-site resource introduces substantial complexity in data harmonization, with heterogeneous collection practices and diverse vocabulary usage patterns requiring sophisticated preprocessing approaches. We present CRISP to unlock the full potential of this valuable resource. CRISP systematically transforms raw Observational Medical Outcomes Partnership Common Data Model data into ML-ready datasets through: (1) transparent data quality management with comprehensive audit trails, (2) cross-vocabulary mapping of heterogeneous medical terminologies to unified SNOMED-CT standards, with deduplication and unit standardization, (3) modular architecture with parallel optimization enabling complete dataset processing in $<$1 day even on standard computing hardware, and (4) comprehensive baseline model benchmarks spanning multiple clinical prediction tasks to establish reproducible performance standards. By providing processing pipeline, baseline implementations, and detailed transformation documentation, CRISP saves researchers months of preprocessing effort and democratizes access to large-scale multi-institutional critical care data, enabling them to focus on advancing clinical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。