arXiv:2412.10804cs.CVcs.AI2024-12AAAI被引 7

为医疗场景构建首个带病征信息的脱敏人脸数据集

Medical Manifestation-Aware De-Identification

  • 用生成与过滤技术重建4万+张真实患者人脸,保护隐私
  • 首次建立医疗场景下带粗/细粒度标注的脱敏基准数据集
  • 融合医学语义先验的简洁方法显著优于已有方案

面部脱敏在通用场景中已有广泛研究,但在医疗场景仍不充分,主要因缺乏大规模患者人脸数据集。本文发布MeMa,包含超过4万张照片级真实感患者人脸,其源自海量真实患者照片,通过精细调控生成与数据筛选流程,在避免泄露真实患者隐私的同时,保留丰富的医学表现特征。我们邀请专家临床医生对MeMa进行粗粒度与细粒度标注,建立了首个医疗场景下的脱敏基准。此外,提出一种基础方法,将数据驱动的医学语义先验融入脱敏过程。尽管方法简洁,但仍显著优于以往方法。数据集可于https://github.com/tianyuan168326/MeMa-Pytorch 获取。

原文摘要 · Abstract (English)

Face de-identification (DeID) has been widely studied for common scenes, but remains under-researched for medical scenes, mostly due to the lack of large-scale patient face datasets. In this paper, we release MeMa, consisting of over 40,000 photo-realistic patient faces. MeMa is re-generated from massive real patient photos. By carefully modulating the generation and data-filtering procedures, MeMa avoids breaching real patient privacy, while ensuring rich and plausible medical manifestations. We recruit expert clinicians to annotate MeMa with both coarse- and fine-grained labels, building the first medical-scene DeID benchmark. Additionally, we propose a baseline approach for this new medical-aware DeID task, by integrating data-driven medical semantic priors into the DeID procedure. Despite its conciseness and simplicity, our approach substantially outperforms previous ones. Dataset is available at https://github.com/tianyuan168326/MeMa-Pytorch.

医疗图像隐私保护生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。