arXiv:2511.02014cs.CV2025-11中稿 · EMBC 2026

用大模型提升医疗影像隐私信息检测,效果因场景而异。

Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images

  • 用大模型做图文联合分析,比传统OCR更准
  • 复杂印刷文字识别准确率提升明显,字错率低至3%-5%
  • 适合有算力但需灵活部署的医疗隐私检测场景

医学影像中保护性健康信息(PHI)的检测对保障患者隐私和符合监管要求至关重要。传统方法主要依赖光学字符识别(OCR)与命名实体识别结合。近年来,大型多模态模型(LMM)为文本提取和语义分析提供了新可能。本研究系统评估了GPT-4o、Gemini 2.5 Flash和Qwen 2.5 7B三款主流闭源与开源LMM,采用两种流程配置:仅文本分析,以及整合OCR与语义分析。结果表明,LMM在OCR性能上显著优于EasyOCR(词错误率WER: 0.03–0.05,字符错误率CER: 0.02–0.03)。然而,此优势并未稳定转化为整体PHI检测准确率提升。在复杂印记模式下表现最优;当文本清晰且对比度高时,不同流程配置结果相近。研究还基于实证提出适配不同条件的LMM选型建议,并设计可扩展、模块化的部署方案。

原文摘要 · Abstract (English)

The detection of Protected Health Information (PHI) in medical imaging is critical for safeguarding patient privacy and ensuring compliance with regulatory frameworks. Traditional detection methodologies predominantly utilize Optical Character Recognition (OCR) models in conjunction with named entity recognition. However, recent advancements in Large Multimodal Model (LMM) present new opportunities for enhanced text extraction and semantic analysis. In this study, we systematically benchmark three prominent closed and open-sourced LMMs, namely GPT-4o, Gemini 2.5 Flash, and Qwen 2.5 7B, utilizing two distinct pipeline configurations: one dedicated to text analysis alone and another integrating both OCR and semantic analysis. Our results indicate that LMM exhibits superior OCR efficacy (WER: 0.03-0.05, CER: 0.02-0.03) compared to conventional models like EasyOCR. However, this improvement in OCR performance does not consistently correlate with enhanced overall PHI detection accuracy. The strongest performance gains are observed on test cases with complex imprint patterns. In scenarios where text regions are well readable with sufficient contrast, and strong LMMs are employed for text analysis after OCR, different pipeline configurations yield similar results. Furthermore, we provide empirically grounded recommendations for LMM selection tailored to specific operational constraints and propose a deployment strategy that leverages scalable and modular infrastructure.

医疗影像大模型隐私检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。