医学影像去标识后仍可还原患者身份,揭示其生物特征属性。
MirrorNet: Can Medical Image Anonymization Really Protect Patient Identity?

- 用双向循环自编码器建立医学影像与人脸图像的对应关系
- 从脱敏影像恢复出可识别的患者肖像(身份区域均方误差0.163)
- 提醒医疗数据应按生物特征严格管理,适合隐私保护研究者
医学影像通常通过去除姓名、日期等元数据进行去标识化,用于研究、教学和公开基准测试,人们普遍认为此举已实现匿名。然而去标识仅移除元数据,未保护像素内容。除了直接包含面部结构的扫描外,影像内容是否仍能识别患者长期未受关注。本文通过一对耦合的循环一致性变分自编码器,学习横断面医学影像与非医学患者身份图像之间的对应关系。实验显示,从独立测试的影像中,模型可重建出可辨识的患者肖像(身份区域均方误差为0.163);反向亦可从人脸图像合成医学影像。结果表明,去标识化的医学影像仍具识别性,实质上相当于患者的影像照片。因此,医学影像数据应被视为生物特征数据而非可匿名记录。为支持复现,代码与训练模型已公开于https://github.com/attilasimko/public-repository。
原文摘要 · Abstract (English)
Medical images are routinely de-identified---names, dates, and other metadata removed---and then shared for research, teaching, and public benchmarks under the assumption that this renders them anonymous. Such de-identification protects the metadata but not the pixels, and---apart from scans that directly contain facial structures---whether the image content itself identifies the patient has received little scrutiny. We investigate this question by learning a cycle-consistent correspondence between a cross-sectional medical image and a non-medical, patient-identifying image, using a pair of coupled, cycle-consistent variational autoencoders. From a held-out scan, the model recovers a recognisable likeness of the patient (identity-region MAE = 0.163); conversely, it synthesises a scan from such an image. These results indicate that a de-identified medical scan remains identifying---it is, in effect, a photograph of the patient---and that imaging data should be governed as biometric data rather than as anonymisable records. To support reproducibility, the code and trained models are shared at https://github.com/attilasimko/public-repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。