arXiv:2504.20220cs.CLcs.CV2025-04被引 4

用视觉语言模型自动提取纸质输血反应单的勾选信息,提升医疗数据录入效率。

A Multimodal Pipeline for Clinical Data Extraction: Applying Vision-Language Models to Scans of Transfusion Reaction Reports

  • 结合检测、多语言OCR与视觉语言模型,自动识别扫描件中的勾选项。
  • 在2017-2024年金标准数据上实现高精度与高召回率。
  • 开源工具可自部署,适合医疗数据数字化场景的开发者使用。

尽管电子健康记录日益普及,许多临床流程仍依赖纸质文档,反映出医疗服务中真实世界的多样性。手动将纸质数据转录为数字格式耗时且易出错。为优化此流程,本研究提出一个开源流水线,用于从扫描文档中提取并分类勾选框数据。以输血反应报告为例,该设计可适配其他包含大量勾选项的文档类型。所提方法融合勾选框检测、多语言光学字符识别(OCR)及多语言视觉语言模型(VLMs)。相比2017至2024年逐年构建的金标准,该流水线表现出高精度与高召回率。结果显著降低行政工作量,并确保监管报告准确性。该流水线开源,支持本地部署,便于对勾选表单进行自主解析。

原文摘要 · Abstract (English)

Despite the growing adoption of electronic health records, many processes still rely on paper documents, reflecting the heterogeneous real-world conditions in which healthcare is delivered. The manual transcription process is time-consuming and prone to errors when transferring paper-based data to digital formats. To streamline this workflow, this study presents an open-source pipeline that extracts and categorizes checkbox data from scanned documents. Demonstrated on transfusion reaction reports, the design supports adaptation to other checkbox-rich document types. The proposed method integrates checkbox detection, multilingual optical character recognition (OCR) and multilingual vision-language models (VLMs). The pipeline achieves high precision and recall compared against annually compiled gold-standards from 2017 to 2024. The result is a reduction in administrative workload and accurate regulatory reporting. The open-source availability of this pipeline encourages self-hosted parsing of checkbox forms.

医疗AI文档解析视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。