基于希伯来语病历构建患者诊疗时间线的模型。
Building Patient Journeys in Hebrew: A Language Model for Clinical Timeline Extraction
- 用五百万份匿名病历持续预训练,适配希伯来语医疗文本。
- 在内外科与肿瘤科数据集上均实现良好时间关系抽取效果。
- 词汇优化提升效率,去标识化不影响性能,适合隐私保护研究。
我们提出一种新的希伯来语医学语言模型,用于从电子健康记录中提取结构化临床时间线,以构建患者诊疗路径。该模型基于 DictaBERT 2.0,基于超过五百万份匿名医院记录进行持续预训练。为评估其有效性,我们构建了两个新数据集——一个来自内科与急诊科,另一个来自肿瘤科,均标注了事件的时间关系。实验结果表明,该模型在两个数据集上均表现优异。此外,我们发现词汇适应可提升词元效率,且去标识化不会影响下游任务性能,支持在保护隐私前提下的模型开发。该模型已开放供研究使用,但需遵守伦理限制。
原文摘要 · Abstract (English)
We present a new Hebrew medical language model designed to extract structured clinical timelines from electronic health records, enabling the construction of patient journeys. Our model is based on DictaBERT 2.0 and continually pre-trained on over five million de-identified hospital records. To evaluate its effectiveness, we introduce two new datasets -- one from internal medicine and emergency departments, and another from oncology -- annotated for event temporal relations. Our results show that our model achieves strong performance on both datasets. We also find that vocabulary adaptation improves token efficiency and that de-identification does not compromise downstream performance, supporting privacy-conscious model development. The model is made available for research use under ethical restrictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。