arXiv:2604.07128cs.CV2026-04

提出隐私保护的医学影像共享方案,兼顾数据可用性与安全。

A Utility-preserving De-identification Pipeline for Cross-hospital Radiology Data Sharing

  • 构建敏感词黑名单与病灶词白名单,生成保留病灶信息的合成图像。
  • 在公开胸片数据集上,去标识后模型诊断准确率接近原始数据。
  • 适合需要跨医院共享影像数据的医疗AI研发团队使用。

大规模放射科数据对构建稳健的医疗AI系统至关重要,但跨医院共享仍受隐私担忧严重制约。现有放射科去标识研究主要关注移除可识别信息以实现合规发布,但去标识数据是否仍能保留足够效用以支持大规模视觉-语言模型训练及跨医院迁移仍未被充分探索。本文提出一种保用型去标识流水线(UPDP),用于跨医院放射科数据共享。具体而言,我们构建了隐私敏感词黑名单和病灶相关词白名单;对于放射科图像,采用生成式过滤机制合成原始图像的隐私过滤且病灶保留版本。这些合成图像与去标识报告可安全跨医院共享,用于下游模型开发与评估。在公开胸片基准测试中,本方法有效移除隐私敏感信息的同时保留了具有诊断意义的病灶线索。在去标识数据上训练的模型保持与原始数据训练模型相当的诊断准确率,同时身份相关准确率显著下降,验证了隐私保护的有效性。在跨医院场景下,进一步证明去标识数据与本地数据结合可带来更优性能。

原文摘要 · Abstract (English)

Large-scale radiology data are critical for developing robust medical AI systems. However, sharing such data across hospitals remains heavily constrained by privacy concerns. Existing de-identification research in radiology mainly focus on removing identifiable information to enable compliant data release. Yet whether de-identified radiology data can still preserve sufficient utility for large-scale vision-language model training and cross-hospital transfer remains underexplored. In this paper, we introduce a utility-preserving de-identification pipeline (UPDP) for cross-hospital radiology data sharing. Specifically, we compile a blacklist of privacy-sensitive terms and a whitelist of pathology-related terms. For radiology images, we use a generative filtering mechanism that synthesis a privacy-filtered and pathology-reserved counterparts of the original images. These synthetic image counterparts, together with ID-filtered reports, can then be securely shared across hospitals for downstream model development and evaluation. Experiments on public chest X-ray benchmarks demonstrate that our method effectively removes privacy-sensitive information while preserving diagnostically relevant pathology cues. Models trained on the de-identified data maintain competitive diagnostic accuracy compared with those trained on the original data, while exhibiting a marked decline in identity-related accuracy, confirming effective privacy protection. In the cross-hospital setting, we further show that de-identified data can be combined with local data to yield better performance.

医学影像隐私保护数据共享生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。