arXiv:2410.09729cs.CVcs.AI2024-10被引 4

用多模态模型从印度手写处方中精准提取药名和剂量

MIRAGE: Multimodal Identification and Recognition of Annotations in Indian General Prescriptions

  • 基于74万张模拟处方图像,微调QWEN VL等多模态大模型
  • 药名与剂量识别准确率达82%,在真实场景下表现稳定
  • 适合医疗信息化、医学图像分析领域的研究者参考

印度医疗机构仍广泛依赖手写病历,尽管电子病历系统已存在,但手写记录增加了统计分析与信息检索的难度。此类记录需专门数据训练模型以识别药物及推荐模式。传统方法使用2D-LSTM,近年有研究尝试用多模态大语言模型(MLLM)进行OCR。本文提出MIRAGE(Multimodal Identification and Recognition of Annotations in Indian General Prescriptions),在743,118张高分辨率模拟处方图像上微调QWEN VL、LLaVA 1.6和Idefics2模型,数据源自印度1,133位医生的全标注处方。实验表明,该方法在提取药名和剂量方面达到82%准确率。

原文摘要 · Abstract (English)

Hospitals in India still rely on handwritten medical records despite the availability of Electronic Medical Records (EMR), complicating statistical analysis and record retrieval. Handwritten records pose a unique challenge, requiring specialized data for training models to recognize medications and their recommendation patterns. While traditional handwriting recognition approaches employ 2-D LSTMs, recent studies have explored using Multimodal Large Language Models (MLLMs) for OCR tasks. Building on this approach, we focus on extracting medication names and dosages from simulated medical records. Our methodology MIRAGE (Multimodal Identification and Recognition of Annotations in indian GEneral prescriptions) involves fine-tuning the QWEN VL, LLaVA 1.6 and Idefics2 models on 743,118 high resolution simulated medical record images-fully annotated from 1,133 doctors across India. Our approach achieves 82% accuracy in extracting medication names and dosages.

多模态OCR医疗文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。