arXiv:2511.04079cs.CL2025-11被引 1

用大规模医学报告训练模型,比商用系统更准地自动隐藏病历隐私信息。

Improving the Performance of Radiology Report De-identification with Large-Scale Training and Benchmarking Against Cloud Vendor Methods

  • 用斯坦福等多源医学报告数据微调Transformer模型,新增年龄类别。
  • 在宾夕法尼亚大学数据集上达到0.973的F1分数,优于所有商用系统。
  • 生成的假隐私信息仍可稳定检测,适合医疗数据安全研究者使用。

目的:通过大规模训练Transformer模型并对比商业云服务,提升放射科报告中受保护健康信息(PHI)的自动化去标识化性能。方法:基于现有先进架构,在斯坦福大学两个大型标注放射科语料库(含胸部X光、胸部CT、腹部/盆腔CT及脑部MRI报告)上进行微调,并引入年龄(AGE)新类别。在斯坦福与宾夕法尼亚大学(Penn)测试集上评估令牌级PHI检测性能,进一步验证了‘显眼藏匿’法生成合成PHI的稳定性,并与商业系统对比。计算各类别下的精确率、召回率和F1分数。结果:模型在宾夕法尼亚大学数据集上总体F1达0.973,在斯坦福数据集上达0.996,优于或持平先前最优模型。合成PHI检测整体F1为0.959(0.958–0.960),在50个独立去标识化宾夕法尼亚数据集上保持一致。该模型在合成宾夕法尼亚报告上表现优于所有厂商系统(整体F1 0.960 vs. 0.632–0.754)。讨论:大规模多模态训练提升了跨机构泛化能力与鲁棒性;合成隐私信息在保障数据可用性的同时确保隐私安全。结论:在多样化放射科数据上训练的Transformer模型,在PHI检测中超越以往学术与商业系统,确立了临床文本安全处理的新基准。

原文摘要 · Abstract (English)

Objective: To enhance automated de-identification of radiology reports by scaling transformer-based models through extensive training datasets and benchmarking performance against commercial cloud vendor systems for protected health information (PHI) detection. Materials and Methods: In this retrospective study, we built upon a state-of-the-art, transformer-based, PHI de-identification pipeline by fine-tuning on two large annotated radiology corpora from Stanford University, encompassing chest X-ray, chest CT, abdomen/pelvis CT, and brain MR reports and introducing an additional PHI category (AGE) into the architecture. Model performance was evaluated on test sets from Stanford and the University of Pennsylvania (Penn) for token-level PHI detection. We further assessed (1) the stability of synthetic PHI generation using a "hide-in-plain-sight" method and (2) performance against commercial systems. Precision, recall, and F1 scores were computed across all PHI categories. Results: Our model achieved overall F1 scores of 0.973 on the Penn dataset and 0.996 on the Stanford dataset, outperforming or maintaining the previous state-of-the-art model performance. Synthetic PHI evaluation showed consistent detectability (overall F1: 0.959 [0.958-0.960]) across 50 independently de-identified Penn datasets. Our model outperformed all vendor systems on synthetic Penn reports (overall F1: 0.960 vs. 0.632-0.754). Discussion: Large-scale, multimodal training improved cross-institutional generalization and robustness. Synthetic PHI generation preserved data utility while ensuring privacy. Conclusion: A transformer-based de-identification model trained on diverse radiology datasets outperforms prior academic and commercial systems in PHI detection and establishes a new benchmark for secure clinical text processing.

医疗文本去标识化Transformer隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。