用AI自动检测医疗图像中的隐私信息,提升数据安全
Exploring AI-based System Design for Pixel-level Protected Health Information Detection in Medical Images
- 分三步走:先定位文本、再提取文字、最后分析内容
- 最佳方案组合使用视觉与语言模型,兼顾速度与成本
- 大模型不仅能识隐私,还能提升识别准确率,适合临床研究
医疗图像去标识化是保障数据共享中隐私安全的关键步骤。该过程的第一步是检测受保护的健康信息(PHI),其可能存在于图像元数据或像素中。尽管此类系统至关重要,但现有基于AI的解决方案缺乏充分评估,制约了可靠工具的发展。本研究提出一个AI驱动的PHI检测流程,包含三个核心模块:文本检测、文本提取与文本分析。我们在两个涵盖多种影像模态和隐私类别的数据集上,对YOLOv11、EasyOCR和GPT-4o三种模型在不同模块配置下的表现进行了基准测试。结果表明,为每个模块分别采用专用视觉与语言模型的组合方案,在性能、延迟和大语言模型使用成本之间实现了良好平衡。此外,大语言模型的应用不仅有助于识别PHI内容,还能增强OCR任务能力,并支持端到端的PHI检测流程,展现出显著成效。
原文摘要 · Abstract (English)
De-identification of medical images is a critical step to ensure privacy during data sharing in research and clinical settings. The initial step in this process involves detecting Protected Health Information (PHI), which can be found in image metadata or imprinted within image pixels. Despite the importance of such systems, there has been limited evaluation of existing AI-based solutions, creating barriers to the development of reliable and robust tools. In this study, we present an AI-based pipeline for PHI detection, comprising three key modules: text detection, text extraction, and text analysis. We benchmark three models - YOLOv11, EasyOCR, and GPT-4o - across different setups corresponding to these modules, evaluating their performance on two different datasets encompassing multiple imaging modalities and PHI categories. Our findings indicate that the optimal setup involves utilizing dedicated vision and language models for each module, which achieves a commendable balance in performance, latency, and cost associated with the usage of Large Language Models (LLMs). Additionally, we show that the application of LLMs not only involves identifying PHI content but also enhances OCR tasks and facilitates an end-to-end PHI detection pipeline, showcasing promising outcomes through our analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。