用新任务头提升Med-BERT预测胰腺癌能力,小样本下效果更优。
Advancing Pancreatic Cancer Prediction with a Next Visit Token Prediction Head on top of Med-BERT
- 将疾病预测转为令牌预测任务,适配Med-BERT预训练目标。
- 在10至500样本下,性能比传统分类高出3%~7%。
- 适合小样本医疗数据场景,尤其适用于罕见病早期预警。
背景:近年来,大量基于大规模数据预训练的通用模型在使用电子健康记录(EHR)进行疾病预测方面表现出色。然而,如何在极小微调数据集下有效利用这些模型仍存在未解问题。方法:我们采用专用于EHR的通用模型Med-BERT,将疾病二分类预测任务重构为令牌预测任务和下一就诊掩码令牌预测任务,使其与Med-BERT的预训练任务格式对齐,以提升胰腺癌(PaCa)在少样本和全监督设置下的预测准确率。结果:将任务重构为令牌预测任务(称为Med-BERT-Sum)在少样本和较大数据样本中表现略优。进一步地,将预测任务重构为下一就诊掩码令牌预测任务(Med-BERT-Mask),在10至500样本的少样本场景下,相比传统二分类任务(Med-BERT-BC)性能提升3%至7%。这些发现表明,使下游任务与模型预训练目标一致,能显著增强模型预测能力,从而提升对罕见与常见疾病的预测效果。结论:重新设计预测任务以匹配通用模型预训练目标,可提高预测准确性,实现更早诊断和及时干预,改善胰腺癌及其他癌症患者的治疗效果、生存率和整体预后。
原文摘要 · Abstract (English)
Background: Recently, numerous foundation models pretrained on extensive data have demonstrated efficacy in disease prediction using Electronic Health Records (EHRs). However, there remains some unanswered questions on how to best utilize such models especially with very small fine-tuning cohorts. Methods: We utilized Med-BERT, an EHR-specific foundation model, and reformulated the disease binary prediction task into a token prediction task and a next visit mask token prediction task to align with Med-BERT's pretraining task format in order to improve the accuracy of pancreatic cancer (PaCa) prediction in both few-shot and fully supervised settings. Results: The reformulation of the task into a token prediction task, referred to as Med-BERT-Sum, demonstrates slightly superior performance in both few-shot scenarios and larger data samples. Furthermore, reformulating the prediction task as a Next Visit Mask Token Prediction task (Med-BERT-Mask) significantly outperforms the conventional Binary Classification (BC) prediction task (Med-BERT-BC) by 3% to 7% in few-shot scenarios with data sizes ranging from 10 to 500 samples. These findings highlight that aligning the downstream task with Med-BERT's pretraining objectives substantially enhances the model's predictive capabilities, thereby improving its effectiveness in predicting both rare and common diseases. Conclusion: Reformatting disease prediction tasks to align with the pretraining of foundation models enhances prediction accuracy, leading to earlier detection and timely intervention. This approach improves treatment effectiveness, survival rates, and overall patient outcomes for PaCa and potentially other cancers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。