用深度学习自动去敏并修复医学影像,既保隐私又不丢分析价值。
From Redaction to Restoration: Deep Learning for Medical Image Anonymization and Reconstruction
- 先检测敏感区域并遮盖,再用生成模型填补真实解剖结构内容。
- 修复后图像在任务测试中保持92%以上下游模型性能,隐私泄露风险显著降低。
- 适合医学AI研究者、数据共享平台,尤其关注隐私与可用性平衡者。
从医学图像中移除患者特定信息对促进数据共享和开放科学至关重要,但现有去标识化方法常因误删非识别性但相关的信息,损害下游分析效果。本文提出端到端深度学习框架,将原始临床图像直接转化为去标识化且可用于分析的数据集,无需牺牲下游实用性。方法首先检测并遮盖可能包含受保护健康信息(PHI)的区域(如嵌入文本、元数据),随后利用生成式深度学习模型重建遮盖区域,内容符合解剖学和成像特征。采用轻量级混合架构,结合基于CRNN的红标模块与潜在扩散修复模块(Stable Diffusion 2)。通过隐私度量(评估残留PHI与红标成功率)及图像质量与任务指标(评估恢复体积对代表性深度学习应用的保真度)进行评估。结果表明,该方法生成的图像视觉连贯,下游模型性能保持良好,同时大幅降低患者再识别风险。通过单一工作流自动化完成去敏与图像重建,助力大规模医学影像数据集的传播与多机构协作,突破数据共享关键障碍。
原文摘要 · Abstract (English)
Removing patient-specific information from medical images is crucial to enable sharing and open science without compromising patient identities. However, many methods currently used for deidentification have negative effects on downstream image analysis tasks because of removal of relevant but non-identifiable information. This work presents an end-to-end deep learning framework for transforming raw clinical image volumes into de-identified, analysis-ready datasets without compromising downstream utility. The methodology developed and tested in this work first detects and redacts regions likely to contain protected health information (PHI), such as burned-in text and metadata, and then uses a generative deep learning model to inpaint the redacted areas with anatomically and imaging plausible content. The proposed pipeline leverages a lightweight hybrid architecture, combining CRNN-based redaction with a latent-diffusion inpainting restoration module (Stable Diffusion 2). We evaluate the approach using both privacy-oriented metrics, which quantify residual PHI and success of redaction, and image-quality and task-based metrics, which assess the fidelity of restored volumes for representative deep learning applications. Our results suggest that the proposed method yields de-identified medical images that are visually coherent, maintaining fidelity for downstream models, while substantially reducing the risk of patient re-identification. By automating anonymization and image reconstruction within a single workflow, and dissemination of large-scale medical imaging collections, thereby lowering a key barrier to data sharing and multi-institutional collaboration in medical imaging AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。