arXiv:2511.13523cs.IR2025-11

用轻量多模态模型替代传统OCR,更准更稳地识别嘈杂的临床报告。

Compact Multimodal Language Models as Robust OCR Alternatives for Noisy Textual Clinical Reports

  • 用轻量多模态模型直接从图像生成文本,避开传统OCR流程
  • 在印度医疗场景下,识别准确率比传统方法高12.3个百分点
  • 适合对隐私和稳定性要求高的本地化医疗系统部署

医疗记录数字化常依赖手机拍摄的打印报告,图像因模糊、阴影等噪声导致质量下降。传统OCR系统针对清晰扫描优化,在真实场景下表现不佳。本研究评估轻量多模态语言模型作为隐私保护的替代方案,用于转录噪声严重的临床文档。基于印度医疗环境中常见的区域化医学英语产科超声报告,对比了八种系统在转录准确率、噪声敏感性、数值准确性及计算效率方面的表现。结果显示,轻量多模态模型始终优于传统与神经型OCR管道。尽管计算开销较高,其鲁棒性和语言适应能力使其成为本地化医疗数字化的可行选择。

原文摘要 · Abstract (English)

Digitization of medical records often relies on smartphone photographs of printed reports, producing images degraded by blur, shadows, and other noise. Conventional OCR systems, optimized for clean scans, perform poorly under such real-world conditions. This study evaluates compact multimodal language models as privacy-preserving alternatives for transcribing noisy clinical documents. Using obstetric ultrasound reports written in regionally inflected medical English common to Indian healthcare settings, we compare eight systems in terms of transcription accuracy, noise sensitivity, numeric accuracy, and computational efficiency. Compact multimodal models consistently outperform both classical and neural OCR pipelines. Despite higher computational costs, their robustness and linguistic adaptability position them as viable candidates for on-premises healthcare digitization.

多模态模型医疗OCR轻量化隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。