arXiv:2507.07599cs.AIcs.CL2025-07被引 2

用微调大模型从急诊分诊记录中自动提取疫苗信息,助力实时安全监测。

Enhancing Vaccine Safety Surveillance: Extracting Vaccine Mentions from Emergency Department Triage Notes Using Fine-Tuned Large Language Models

  • 微调Llama 3模型识别急诊记录中的疫苗名称
  • 30亿参数模型准确率优于提示工程与规则方法
  • 量化后可在资源受限环境部署,适合疾控系统使用

本研究评估微调后的Llama 3.2模型在急诊分诊记录中提取疫苗相关信息的能力,以支持近实时疫苗安全监测。通过提示工程生成初始标注数据,并由人工确认。比较了提示工程模型、微调模型和基于规则的方法的性能。30亿参数的微调版Llama 3模型在提取疫苗名称方面表现最佳。模型量化技术实现了在资源受限环境下的高效部署。结果表明,大语言模型可有效自动化从急诊记录中提取数据,支持疫苗安全监测并实现免疫接种后不良事件的早期发现。

原文摘要 · Abstract (English)

This study evaluates fine-tuned Llama 3.2 models for extracting vaccine-related information from emergency department triage notes to support near real-time vaccine safety surveillance. Prompt engineering was used to initially create a labeled dataset, which was then confirmed by human annotators. The performance of prompt-engineered models, fine-tuned models, and a rule-based approach was compared. The fine-tuned Llama 3 billion parameter model outperformed other models in its accuracy of extracting vaccine names. Model quantization enabled efficient deployment in resource-constrained environments. Findings demonstrate the potential of large language models in automating data extraction from emergency department notes, supporting efficient vaccine safety surveillance and early detection of emerging adverse events following immunization issues.

疫苗监测大模型应用医疗文本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。