arXiv:2507.21703cs.CV2025-07ICCV被引 3

医学影像去标识化新方法,兼顾诊断语义与可调隐私等级

Semantics versus Identity: A Divide-and-Conquer Approach towards Adjustable Medical Image De-Identification

  • 分步处理:先阻断身份区域,再用医疗大模型补全语义
  • 在7个数据集上实现最优去标识效果,下游任务性能领先
  • 适合需要灵活控制隐私强度的医疗AI应用

医学影像推动了辅助诊断发展,但重识别风险带来严重隐私问题,亟需有效的去标识技术。现有方法无法同时兼顾医疗语义保留与隐私等级可调性。为此,本文提出一种分而治之框架:第一步为身份阻断,通过遮蔽不同比例的身份相关区域实现多级隐私保护;第二步为医学语义补偿,利用预训练医学基础模型(MFMs)提取医疗语义特征,恢复被遮挡区域的信息。此外,考虑到MFMs特征仍可能包含残留身份信息,我们引入基于最小描述长度原则的特征解耦策略,有效分离并丢弃身份成分。在七个数据集和三个下游任务上的广泛评估表明,该方法达到当前最优性能。

原文摘要 · Abstract (English)

Medical imaging has significantly advanced computer-aided diagnosis, yet its re-identification (ReID) risks raise critical privacy concerns, calling for de-identification (DeID) techniques. Unfortunately, existing DeID methods neither particularly preserve medical semantics, nor are flexibly adjustable towards different privacy levels. To address these issues, we propose a divide-and-conquer framework comprising two steps: (1) Identity-Blocking, which blocks varying proportions of identity-related regions, to achieve different privacy levels; and (2) Medical-Semantics-Compensation, which leverages pre-trained Medical Foundation Models (MFMs) to extract medical semantic features to compensate the blocked regions. Moreover, recognizing that features from MFMs may still contain residual identity information, we introduce a Minimum Description Length principle-based feature decoupling strategy, to effectively decouple and discard such identity components. Extensive evaluations against existing approaches across seven datasets and three downstream tasks, demonstrates our state-of-the-art performance.

医学图像去标识化隐私保护大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。