arXiv:2502.07516eess.IVcs.AI2025-02被引 4

去标识化痕迹反而加剧了医学图像生成中的数据记忆风险。

The Devil is in the Prompts: De-Identification Traces Enhance Memorization Risks in Synthetic Chest X-Ray Generation

  • 分析发现去标识化标记文本最易被模型记忆
  • 去标识化标记贡献了最高比例的记忆化内容
  • 现有隐私保护方法无效,需改进策略

生成模型,尤其是文本到图像(T2I)扩散模型,在医学图像分析中起关键作用。然而,这些模型容易对训练数据产生记忆,严重威胁患者隐私。合成胸部X光片是医学图像分析中最常见的应用之一,主要依赖于MIMIC-CXR数据集。本研究首次系统性识别出在MIMIC-CXR中导致训练数据记忆的关键提示词和文本标记。分析揭示两个意外发现:(1)包含去标识化过程痕迹(用于隐藏受保护健康信息的标记)的提示词最易被记忆;(2)在所有文本标记中,去标识化标记对记忆化的贡献最大。这暴露了标准匿名化实践与使用MIMIC-CXR进行T2I合成中存在的更广泛问题。更严重的是,现有的推理阶段记忆缓解策略无效,无法充分降低模型对记忆文本标记的依赖。为此,我们提出了针对不同利益相关方的可操作策略,以提升医疗影像生成模型的隐私保护能力和可靠性。最后,本研究为未来基于MIMIC-CXR开发和评估记忆缓解技术奠定了基础。代码已匿名公开于https://anonymous.4open.science/r/diffusion_memorization-8011/

原文摘要 · Abstract (English)

Generative models, particularly text-to-image (T2I) diffusion models, play a crucial role in medical image analysis. However, these models are prone to training data memorization, posing significant risks to patient privacy. Synthetic chest X-ray generation is one of the most common applications in medical image analysis with the MIMIC-CXR dataset serving as the primary data repository for this task. This study presents the first systematic attempt to identify prompts and text tokens in MIMIC-CXR that contribute the most to training data memorization. Our analysis reveals two unexpected findings: (1) prompts containing traces of de-identification procedures (markers introduced to hide Protected Health Information) are the most memorized, and (2) among all tokens, de-identification markers contribute the most towards memorization. This highlights a broader issue with the standard anonymization practices and T2I synthesis with MIMIC-CXR. To exacerbate, existing inference-time memorization mitigation strategies are ineffective and fail to sufficiently reduce the model's reliance on memorized text tokens. On this front, we propose actionable strategies for different stakeholders to enhance privacy and improve the reliability of generative models in medical imaging. Finally, our results provide a foundation for future work on developing and benchmarking memorization mitigation techniques for synthetic chest X-ray generation using the MIMIC-CXR dataset. The anonymized code is available at https://anonymous.4open.science/r/diffusion_memorization-8011/

医学图像隐私安全扩散模型数据记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。