arXiv:2409.09501cs.CLcs.AI2024-09被引 7

用合成数据生成临床病历,解决隐私问题并支持医疗研究。

Synthetic4Health: Generating Annotated Synthetic Clinical Letters

  • 基于预训练模型与掩码策略生成去标识化病历
  • 保留临床实体和文档结构的合成文本更贴近真实数据
  • 适合医疗AI研究、教学及数据扩展场景

由于临床病历包含敏感信息,相关数据集难以广泛用于模型训练、医学研究和教学。本文旨在生成可靠、多样且去标识化的合成临床病历。通过探索不同预训练语言模型(PLMs)进行文本掩码与生成,重点实验了基于生物医学临床语料训练的Bio_ClinicalBERT模型,并比较了多种掩码策略。采用定性与定量方法评估生成效果,同时通过命名实体识别(NER)下游任务检验其可用性。结果表明:1)编码器类模型优于编码器-解码器模型;2)通用语料训练的编码器模型在保留临床信息时表现可媲美专有临床数据训练模型;3)保留临床实体与文档结构比单纯微调模型更符合目标;4)掩码策略影响显著:掩码停用词提升质量,掩码名词或动词则降低质量;5)评价中应以BERTScore为主,其他指标为辅;6)上下文信息对模型理解影响不大,合成病历具备替代原始数据开展下游任务的潜力。

原文摘要 · Abstract (English)

Since clinical letters contain sensitive information, clinical-related datasets can not be widely applied in model training, medical research, and teaching. This work aims to generate reliable, various, and de-identified synthetic clinical letters. To achieve this goal, we explored different pre-trained language models (PLMs) for masking and generating text. After that, we worked on Bio\_ClinicalBERT, a high-performing model, and experimented with different masking strategies. Both qualitative and quantitative methods were used for evaluation. Additionally, a downstream task, Named Entity Recognition (NER), was also implemented to assess the usability of these synthetic letters. The results indicate that 1) encoder-only models outperform encoder-decoder models. 2) Among encoder-only models, those trained on general corpora perform comparably to those trained on clinical data when clinical information is preserved. 3) Additionally, preserving clinical entities and document structure better aligns with our objectives than simply fine-tuning the model. 4) Furthermore, different masking strategies can impact the quality of synthetic clinical letters. Masking stopwords has a positive impact, while masking nouns or verbs has a negative effect. 5) For evaluation, BERTScore should be the primary quantitative evaluation metric, with other metrics serving as supplementary references. 6) Contextual information does not significantly impact the models' understanding, so the synthetic clinical letters have the potential to replace the original ones in downstream tasks.

临床生成数据合成隐私保护NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。