arXiv:2507.23736stat.MLcs.LG2025-07被引 6

混合AI与规则的去标识化框架,实现医疗影像数据安全共享。

DICOM De-Identification via Hybrid AI and Rule-Based Framework for Scalable, Uncertainty-Aware Redaction

  • 结合规则与大模型识别元数据和像素中的敏感信息
  • 通过不确定性量化提升检测可靠性,降低漏检风险
  • 符合HIPAA等标准,适合医学研究数据合规发布

医疗影像及文本数据的开放可推动医疗研究进步,但DICOM文件中包含的受保护健康信息(PHI)和个人身份信息(PII)成为数据共享的主要障碍。本文提出由Impact Business Information Solutions(IBIS)开发的混合去标识化框架,融合规则与AI技术,并引入严格的不确定性量化机制,全面清除元数据与像素数据中的敏感信息。该方法首先采用双层规则系统处理显式与推断型元数据元素,进一步结合微调的大语言模型(LLM)进行命名实体识别(NER),训练数据为模拟真实临床场景的合成数据集。对于像素数据,采用不确定性感知的Faster R-CNN定位嵌入文本,通过OCR提取候选信息,并经NER流程完成最终去标识化。关键的是,不确定性量化提供置信度评估,增强自动化可靠性,支持人工介入验证以控制残留风险。该框架在多项基准数据集及监管标准(包括DICOM、HIPAA、TCIA)上表现稳健。通过规模化自动化、不确定性量化与严格质控,本方案有效应对医疗数据去标识化的关键挑战,支持科研数据的安全、伦理与可信释放。

原文摘要 · Abstract (English)

Access to medical imaging and associated text data has the potential to drive major advances in healthcare research and patient outcomes. However, the presence of Protected Health Information (PHI) and Personally Identifiable Information (PII) in Digital Imaging and Communications in Medicine (DICOM) files presents a significant barrier to the ethical and secure sharing of imaging datasets. This paper presents a hybrid de-identification framework developed by Impact Business Information Solutions (IBIS) that combines rule-based and AI-driven techniques, and rigorous uncertainty quantification for comprehensive PHI/PII removal from both metadata and pixel data. Our approach begins with a two-tiered rule-based system targeting explicit and inferred metadata elements, further augmented by a large language model (LLM) fine-tuned for Named Entity Recognition (NER), and trained on a suite of synthetic datasets simulating realistic clinical PHI/PII. For pixel data, we employ an uncertainty-aware Faster R-CNN model to localize embedded text, extract candidate PHI via Optical Character Recognition (OCR), and apply the NER pipeline for final redaction. Crucially, uncertainty quantification provides confidence measures for AI-based detections to enhance automation reliability and enable informed human-in-the-loop verification to manage residual risks. This uncertainty-aware deidentification framework achieves robust performance across benchmark datasets and regulatory standards, including DICOM, HIPAA, and TCIA compliance metrics. By combining scalable automation, uncertainty quantification, and rigorous quality assurance, our solution addresses critical challenges in medical data de-identification and supports the secure, ethical, and trustworthy release of imaging data for research.

医疗影像去标识化AI安全不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。