arXiv:2504.20547cs.CL2025-04被引 9

用大模型处理电子病历,文本方法表现不输传统表格模型

Revisiting the MIMIC-IV Benchmark: Experiments Using Language Models for Electronic Health Records

  • 将病历数据转为文本,用模板构建输入
  • 微调后的大模型在死亡预测上媲美专业表格分类器
  • 零样本模型难有效利用病历信息,适合研究者参考

医疗领域缺乏标准的文本评估基准,制约了自然语言模型在健康相关任务中的应用。本文重新审视公开的MIMIC-IV电子病历(EHR)基准,首先将该数据集集成至Hugging Face数据集库,便于共享使用;其次探索使用模板将表格型EHR数据转化为文本形式。在患者死亡率预测任务上,微调后的文本模型表现与强健的表格分类器相当,而零样本大模型则难以有效利用病历表征。本研究展示了文本方法在医疗领域的潜力,并指出了未来改进方向。

原文摘要 · Abstract (English)

The lack of standardized evaluation benchmarks in the medical domain for text inputs can be a barrier to widely adopting and leveraging the potential of natural language models for health-related downstream tasks. This paper revisited an openly available MIMIC-IV benchmark for electronic health records (EHRs) to address this issue. First, we integrate the MIMIC-IV data within the Hugging Face datasets library to allow an easy share and use of this collection. Second, we investigate the application of templates to convert EHR tabular data to text. Experiments using fine-tuned and zero-shot LLMs on the mortality of patients task show that fine-tuned text-based models are competitive against robust tabular classifiers. In contrast, zero-shot LLMs struggle to leverage EHR representations. This study underlines the potential of text-based approaches in the medical field and highlights areas for further improvement.

电子病历大模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。