arXiv:2506.12808cs.CV2025-06综述被引 11

梳理MIMIC数据集在数字健康中的挑战与进展,助力医疗AI落地

Leveraging MIMIC Datasets for Better Digital Health: A Review on Open Problems, Progress Highlights, and Future Promises

  • 系统分析MIMIC数据在粒度、编码、伦理等方面的瓶颈
  • 指出降维、时序建模、因果推断等关键突破方向
  • 适合关注医疗数据应用的科研人员与临床工程师

医学信息集市(MIMIC)数据集已成为数字健康研究的核心资源,提供数万例重症监护入院的匿名记录,广泛应用于临床决策支持、预后预测和医疗数据分析。尽管已有大量研究探讨基于MIMIC模型的预测能力与临床价值,但数据整合、表征与互操作性等关键挑战仍缺乏深入探讨。本文聚焦开放问题,识别出数据粒度粗糙、基数受限、编码体系异构及伦理约束等阻碍模型泛化与实时部署的问题。同时,总结了在降维、时序建模、因果推断与隐私保护分析方面的进展,并展望了混合建模、联邦学习与标准化预处理流程等未来方向。通过剖析结构性局限及其影响,本综述为下一代基于MIMIC的数字健康创新提供了可操作的洞察。

原文摘要 · Abstract (English)

The Medical Information Mart for Intensive Care (MIMIC) datasets have become the Kernel of Digital Health Research by providing freely accessible, deidentified records from tens of thousands of critical care admissions, enabling a broad spectrum of applications in clinical decision support, outcome prediction, and healthcare analytics. Although numerous studies and surveys have explored the predictive power and clinical utility of MIMIC based models, critical challenges in data integration, representation, and interoperability remain underexplored. This paper presents a comprehensive survey that focuses uniquely on open problems. We identify persistent issues such as data granularity, cardinality limitations, heterogeneous coding schemes, and ethical constraints that hinder the generalizability and real-time implementation of machine learning models. We highlight key progress in dimensionality reduction, temporal modelling, causal inference, and privacy preserving analytics, while also outlining promising directions including hybrid modelling, federated learning, and standardized preprocessing pipelines. By critically examining these structural limitations and their implications, this survey offers actionable insights to guide the next generation of MIMIC powered digital health innovations.

医疗AI数据集数字健康综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。