arXiv:2411.18940cs.CL2024-11被引 5

用大模型重写病历生成合成数据,助力医疗语言模型预训练

Rephrasing Electronic Health Records for Pretraining Clinical Language Models

  • 用4个小型大模型重写真实病历生成合成文本
  • 合成数据提升语言建模与下游任务性能,优于以往方法
  • 小规模增量即可增效,适合机构级或大规模应用

临床语言模型在医疗应用中至关重要,但其训练依赖大量临床文本。由于患者隐私问题,从电子健康记录(EHR)中大规模获取临床笔记面临挑战。本研究利用大语言模型(LLMs)重写现有临床笔记,生成合成预训练语料,借鉴此前对网络数据重写的思路。我们测试了四种小型模型(<10B参数)生成合成临床文本,并用于预训练解码器型与编码器型语言模型。结果表明,该方法在语言建模和下游任务中表现优于无需参考真实临床文本的先前合成方法。将原始临床笔记与来自不同大模型的合成语料结合,即使在小令牌预算下也能提升性能,展示了该方法在机构级预训练或大规模合成中的潜力。

原文摘要 · Abstract (English)

Clinical language models are important for many applications in healthcare, but their development depends on access to extensive clinical text for pretraining. However, obtaining clinical notes from electronic health records (EHRs) at scale is challenging due to patient privacy concerns. In this study, we rephrase existing clinical notes using LLMs to generate synthetic pretraining corpora, drawing inspiration from previous work on rephrasing web data. We examine four popular small-sized LLMs (<10B) to create synthetic clinical text to pretrain both decoder-based and encoder-based language models. The method yields better results in language modeling and downstream tasks than previous synthesis approaches without referencing real clinical text. We find that augmenting original clinical notes with synthetic corpora from different LLMs improves performances even at a small token budget, showing the potential of this method to support pretraining at the institutional level or be scaled to synthesize large-scale clinical corpora.

医疗AI大模型数据合成自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。