arXiv:2409.09831cs.CLcs.LG2024-09NAACL被引 5

用掩码语言模型生成低重识别风险的合成病历,兼顾隐私与数据质量。

Generating Synthetic Free-text Medical Records with Low Re-identification Risk using Masked Language Modeling

  • 基于掩码语言模型生成病历,控制多样性并降低重识别风险。
  • 合成数据符合HIPAA标准,敏感信息召回率达96%,重识别风险仅3.5%。
  • 模型仅120M参数,推理成本低,适合实际医疗数据应用。

大量可用的医疗记录具有提升医疗和生物医学研究的潜力,但受隐私限制,仅限内部使用。现有方法多采用因果语言建模生成合成数据,但难以在保证患者隐私的同时控制生成多样性,且生成成本较高。本文提出一种基于掩码语言模型的合成自由文本病历生成系统,可在保留关键医疗信息的同时引入生成多样性,并显著降低重识别风险。系统规模约为120M参数,推理开销小。实验结果表明,生成数据具备高质量,敏感信息(PHI)召回率达到96%,重识别风险仅为3.5%。下游任务评估显示,该合成数据可有效训练模型,性能接近真实数据。

原文摘要 · Abstract (English)

The vast amount of available medical records has the potential to improve healthcare and biomedical research. However, privacy restrictions make these data accessible for internal use only. Recent works have addressed this problem by generating synthetic data using Causal Language Modeling. Unfortunately, by taking this approach, it is often impossible to guarantee patient privacy while offering the ability to control the diversity of generations without increasing the cost of generating such data. In contrast, we present a system for generating synthetic free-text medical records using Masked Language Modeling. The system preserves critical medical information while introducing diversity in the generations and minimising re-identification risk. The system's size is about 120M parameters, minimising inference cost. The results demonstrate high-quality synthetic data with a HIPAA-compliant PHI recall rate of 96% and a re-identification risk of 3.5%. Moreover, downstream evaluations show that the generated data can effectively train a model with performance comparable to real data.

合成数据医疗生成隐私保护语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。