针对电子病历标签噪声,提出动态校准与增强框架,提升预测准确性。
Dynamical Label Augmentation and Calibration for Noisy Electronic Health Records
- 基于贝塔混合模型区分确定与不确定样本,动态调整标签。
- 在eICU和MIMIC-IV-ED数据集上,高噪声下仍保持领先性能。
- 适合处理带标签噪声的医疗时间序列预测任务。
医学研究,尤其是患者预后预测,严重依赖从电子健康记录(EHR)中提取的医疗时间序列数据,这些数据提供了丰富的患者病史信息。尽管经过严格审查,标签错误仍不可避免,且会显著影响患者结局预测的准确性。为应对这一挑战,我们提出一种基于注意力的学习框架——针对时间序列标签噪声学习的动态校准与增强框架(ACTLL)。该框架利用双成分贝塔混合模型,根据各类别拟合度分布识别出确定和不确定样本集,并在捕捉全局时间动态的同时,对不确定样本进行动态标签校准,或从确定样本集中增强可信实例。在大规模EHR数据集eICU、MIMIC-IV-ED以及来自UCR和UEA存储库的多个基准数据集上的实验结果表明,本模型在高噪声条件下实现了最先进的性能。
原文摘要 · Abstract (English)
Medical research, particularly in predicting patient outcomes, heavily relies on medical time series data extracted from Electronic Health Records (EHR), which provide extensive information on patient histories. Despite rigorous examination, labeling errors are inevitable and can significantly impede accurate predictions of patient outcome. To address this challenge, we propose an \textbf{A}ttention-based Learning Framework with Dynamic \textbf{C}alibration and Augmentation for \textbf{T}ime series Noisy \textbf{L}abel \textbf{L}earning (ACTLL). This framework leverages a two-component Beta mixture model to identify the certain and uncertain sets of instances based on the fitness distribution of each class, and it captures global temporal dynamics while dynamically calibrating labels from the uncertain set or augmenting confident instances from the certain set. Experimental results on large-scale EHR datasets eICU and MIMIC-IV-ED, and several benchmark datasets from the UCR and UEA repositories, demonstrate that our model ACTLL has achieved state-of-the-art performance, especially under high noise levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。