arXiv:2504.16448cs.CLcs.AI2025-04被引 1

用轻量微调+代码提示,把医生问诊对话转成结构化病历。

EMRModel: A Large Language Model for Extracting Medical Consultation Dialogues into Structured Medical Records

  • 结合LoRA微调与代码风格提示,提升病历结构化提取效率。
  • 在真实对话数据上达到88.1%的F1值,比通用模型高49.5%。
  • 专为医疗对话设计评估基准,适合临床NLP研究者参考。

医生问诊对话包含关键临床信息,但其非结构化特性限制了在诊疗中的有效利用。传统基于规则或浅层机器学习的方法难以捕捉深层隐含语义。近期,大预训练语言模型与低秩适应(LoRA)这一轻量级微调方法展现出在结构化信息抽取中的潜力。我们提出EMRModel,通过融合LoRA微调与代码风格提示设计,旨在高效将医疗问诊对话转化为结构化电子病历(EMRs)。此外,我们构建了一个高质量、真实场景下的问诊对话数据集,并附有详细标注。同时,我们引入细粒度的问诊信息抽取评估基准,提供系统化评估方法,推动医疗自然语言处理模型优化。实验表明,EMRModel取得88.1%的F1分数,相比标准预训练模型提升49.5%。相较于传统LoRA微调方法,本模型表现更优,验证了其在结构化病历提取任务中的有效性。

原文摘要 · Abstract (English)

Medical consultation dialogues contain critical clinical information, yet their unstructured nature hinders effective utilization in diagnosis and treatment. Traditional methods, relying on rule-based or shallow machine learning techniques, struggle to capture deep and implicit semantics. Recently, large pre-trained language models and Low-Rank Adaptation (LoRA), a lightweight fine-tuning method, have shown promise for structured information extraction. We propose EMRModel, a novel approach that integrates LoRA-based fine-tuning with code-style prompt design, aiming to efficiently convert medical consultation dialogues into structured electronic medical records (EMRs). Additionally, we construct a high-quality, realistically grounded dataset of medical consultation dialogues with detailed annotations. Furthermore, we introduce a fine-grained evaluation benchmark for medical consultation information extraction and provide a systematic evaluation methodology, advancing the optimization of medical natural language processing (NLP) models. Experimental results show EMRModel achieves an F1 score of 88.1%, improving by49.5% over standard pre-trained models. Compared to traditional LoRA fine-tuning methods, our model shows superior performance, highlighting its effectiveness in structured medical record extraction tasks.

医疗NLP结构化大模型对话提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。