用真实伪数据提升医疗文本去标识化模型性能。
Enhancing Clinical Models with Pseudo Data for De-identification
- 用真实伪数据替换脱敏文本,训练更有效的临床模型。
- 在去标识化任务中显著优于现有基线方法。
- 开源伪数据生成工具与模型代码,适合医疗AI研究者。
许多模型因隐私原因在去标识化文本上预训练。临床基础模型通常在使用掩码语法替代受保护健康信息的去标识化文本上训练。尽管这些模型日益流行,但对其在脱敏文本上训练效果的研究仍较少。本文在包含脱敏文本及替换为真实伪文本的版本的数据集上,预训练多个仅编码器模型,并微调用于受保护健康信息去标识化任务,结果表明我们的方法显著优于先前基线。贡献包括:a)训练建议的新发现,出人意料;b)用于生成伪数据的脱敏文本替换策略;c)预训练嵌入和微调后的特定任务模型;d)实验中使用的伪数据生成与模型源代码完全开源。
原文摘要 · Abstract (English)
Many models are pretrained on redacted text for privacy reasons. Clinical foundation models are often trained on de-identified text, which uses special syntax (masked) text in place of protected health information. Even though these models have increased in popularity, there has been little effort in understanding the effects of training them on redacted text. In this work, we pretrain several encoder-only models on a dataset that contains redacted text and a version with replaced realistic pseudo text. We then fine-tuned models for the protected health information de-identification task and show how our methods significantly outperform previous baselines. The contributions of this work include: a) our novel, and yet surprising findings with training recommendations, b) redacted text replacements used to produce the pseudo dataset, c) pretrained embeddings and fine-tuned task specific models, and d) freely available pseudo training dataset generation and model source code used in our experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。