arXiv:2501.14105cs.CLcs.AI2025-01被引 3

用开源大模型实现更安全高效的病历分段,效果超越商用模型。

MedSlice: Fine-Tuned Large Language Models for Secure Clinical Note Sectioning

  • 微调开源大模型处理病历三类关键段落
  • 在外部测试集上仍达F1=0.85,优于GPT-4o
  • 适合关注隐私与成本的医疗AI研发团队

从临床病历中提取章节对下游分析至关重要,但因格式差异和人工标注耗时而困难。尽管专有大语言模型(LLMs)表现良好,但隐私顾虑限制其使用。本研究开发了基于开源LLM的自动化病历分段流程,聚焦主诉、间隔史和评估计划三类内容。使用487份病历构建的标注数据集,对三个开源LLM进行微调,并与专有模型(GPT-4o、GPT-4o mini)对比性能,评估指标为精确率、召回率与F1分数。微调后的Llama 3.1 8B模型在内部测试中取得F1=0.92,超过GPT-4o;在外部验证集上仍保持高精度(F1=0.85)。结果表明,经微调的开源模型在临床病历分段任务中可超越专有模型,在成本、性能与可访问性方面具有显著优势。

原文摘要 · Abstract (English)

Extracting sections from clinical notes is crucial for downstream analysis but is challenging due to variability in formatting and labor-intensive nature of manual sectioning. While proprietary large language models (LLMs) have shown promise, privacy concerns limit their accessibility. This study develops a pipeline for automated note sectioning using open-source LLMs, focusing on three sections: History of Present Illness, Interval History, and Assessment and Plan. We fine-tuned three open-source LLMs to extract sections using a curated dataset of 487 progress notes, comparing results relative to proprietary models (GPT-4o, GPT-4o mini). Internal and external validity were assessed via precision, recall and F1 score. Fine-tuned Llama 3.1 8B outperformed GPT-4o (F1=0.92). On the external validity test set, performance remained high (F1= 0.85). Fine-tuned open-source LLMs can surpass proprietary models in clinical note sectioning, offering advantages in cost, performance, and accessibility.

医疗AI大模型微调病历分段开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。