构建合成医疗影像数据集与评估工具,提升去标识化可靠性
Medical Image De-Identification Resources: Synthetic DICOM Data and Tools for Validation
- 用真实数据生成含伪造隐私信息的合成DICOM数据
- 涵盖538名患者、5万余张图像,支持自动化评估去标识效果
- 符合HIPAA标准,适合研究者和机构验证数据共享安全性
医学影像研究依赖大规模数据共享以提升可复现性并训练人工智能模型,但患者隐私保护仍是开放共享的主要挑战。数字成像与通信(DICOM)作为全球医学影像标准格式,包含关键临床元数据及大量受保护健康信息(PHI)和个人身份信息(PII)。有效的去标识化需移除标识符、保留科学价值并维持DICOM有效性。现有工具多缺乏客观评估手段,依赖主观审查,影响可复现性与监管信心。为此,我们构建了公开可用的合成隐私信息嵌入式DICOM数据集——医学影像去标识化(MIDI)数据集,基于来自癌症影像档案馆(TCIA)的公开去标识数据生成。数据集包含538名受试者(216用于验证,322用于测试)、605项研究、708个序列及53,581个DICOM图像实例,覆盖多个厂商、成像模态和癌种类型。通过在结构化数据元素、纯文本元素和像素数据中嵌入合成的PHI/PII,模拟TCIA归档团队实际遇到的身份泄露场景。配套评估工具包括Python脚本、答案密钥(已知真值)和映射文件,支持对清洗后数据与预期转换进行自动化比对。该框架遵循HIPAA隐私规则“安全港”方法、DICOM PS3.15保密性配置文件及TCIA最佳实践,支持客观、标准化的去标识化工作流评估,推动更安全、一致的医学影像共享。
原文摘要 · Abstract (English)
Medical imaging research increasingly depends on large-scale data sharing to promote reproducibility and train Artificial Intelligence (AI) models. Ensuring patient privacy remains a significant challenge for open-access data sharing. Digital Imaging and Communications in Medicine (DICOM), the global standard data format for medical imaging, encodes both essential clinical metadata and extensive protected health information (PHI) and personally identifiable information (PII). Effective de-identification must remove identifiers, preserve scientific utility, and maintain DICOM validity. Tools exist to perform de-identification, but few assess its effectiveness, and most rely on subjective reviews, limiting reproducibility and regulatory confidence. To address this gap, we developed an openly accessible DICOM dataset infused with synthetic PHI/PII and an evaluation framework for benchmarking image de-identification workflows. The Medical Image de-identification (MIDI) dataset was built using publicly available de-identified data from The Cancer Imaging Archive (TCIA). It includes 538 subjects (216 for validation, 322 for testing), 605 studies, 708 series, and 53,581 DICOM image instances. These span multiple vendors, imaging modalities, and cancer types. Synthetic PHI and PII were embedded into structured data elements, plain text data elements, and pixel data to simulate real-world identity leaks encountered by TCIA curation teams. Accompanying evaluation tools include a Python script, answer keys (known truth), and mapping files that enable automated comparison of curated data against expected transformations. The framework is aligned with the HIPAA Privacy Rule "Safe Harbor" method, DICOM PS3.15 Confidentiality Profiles, and TCIA best practices. It supports objective, standards-driven evaluation of de-identification workflows, promoting safer and more consistent medical image sharing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。