评测医疗影像去标识化工具,确保隐私安全同时保留研究有用信息。
Medical Image De-Identification Benchmark Challenge
- 基于HIPAA标准构建真实影像数据集,人工注入假标识信息用于测试。
- 97.91%至99.93%的去标识准确率,十支团队完成最终测试。
- 适合医学影像隐私保护、AI研究数据共享领域从业者参考。
医疗影像去标识化(deID)是共享医学影像以符合患者隐私法规的基本要求,尤其在公开数据库中尤为重要。同时,保留非标识性元数据对支持医学影像人工智能(AI)的下游研究也至关重要。MIDI-B挑战赛旨在建立一个标准化平台,用于评估符合HIPAA安全港规定、DICOM属性保密配置文件及癌症影像档案(TCIA)定义的最佳实践的DICOM图像去标识工具。挑战赛采用大规模、多样化、多中心、多模态的真实去标识影像,人工插入合成的受保护健康信息(PHI)和可识别信息(PII)。MIDI-B分为训练、验证和测试三个阶段。共有80人注册,10支团队成功完成测试阶段。测试阶段使用包含216和322名受试者合成标识的DICOM图像。通过计算正确操作占总需操作的比例评估表现,得分范围为97.91%至99.93%。参与者使用开源与专有工具、定制配置、大语言模型及光学字符识别(OCR)技术。本文全面报告了挑战赛的设计、实施、结果与经验教训。
原文摘要 · Abstract (English)
The de-identification (deID) of protected health information (PHI) and personally identifiable information (PII) is a fundamental requirement for sharing medical images, particularly through public repositories, to ensure compliance with patient privacy laws. In addition, preservation of non-PHI metadata to inform and enable downstream development of imaging artificial intelligence (AI) is an important consideration in biomedical research. The goal of MIDI-B was to provide a standardized platform for benchmarking of DICOM image deID tools based on a set of rules conformant to the HIPAA Safe Harbor regulation, the DICOM Attribute Confidentiality Profiles, and best practices in preservation of research-critical metadata, as defined by The Cancer Imaging Archive (TCIA). The challenge employed a large, diverse, multi-center, and multi-modality set of real de-identified radiology images with synthetic PHI/PII inserted. The MIDI-B Challenge consisted of three phases: training, validation, and test. Eighty individuals registered for the challenge. In the training phase, we encouraged participants to tune their algorithms using their in-house or public data. The validation and test phases utilized the DICOM images containing synthetic identifiers (of 216 and 322 subjects, respectively). Ten teams successfully completed the test phase of the challenge. To measure success of a rule-based approach to image deID, scores were computed as the percentage of correct actions from the total number of required actions. The scores ranged from 97.91% to 99.93%. Participants employed a variety of open-source and proprietary tools with customized configurations, large language models, and optical character recognition (OCR). In this paper we provide a comprehensive report on the MIDI-B Challenge's design, implementation, results, and lessons learned.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。