arXiv:2409.10576cs.CLcs.IR2024-09被引 2

用大模型和检索增强生成,自动从病历报告中提取结构化数据。

Language Models and Retrieval Augmented Generation for Automated Structured Data Extraction from Diagnostic Reports

  • 基于大语言模型与检索增强生成,构建自动化提取流水线。
  • 对脑肿瘤报告提取准确率超98%,病理报告超90%。
  • 提示工程与领域微调显著提升效果,适合医疗研究场景。

目的:开发并评估一种基于开源大语言模型(LMs)与检索增强生成(RAG)的自动化系统,用于从非结构化放射科和病理科报告中提取结构化临床信息,并分析模型配置变量对提取性能的影响。方法与材料:研究使用两个数据集——7,294份标注了脑肿瘤报告与数据系统(BT-RADS)评分的放射科报告,以及2,154份标注了异柠檬酸脱氢酶(IDH)突变状态的病理科报告。构建自动化流程,对比多种大语言模型及RAG配置的性能。系统评估了模型规模、量化方式、提示策略、输出格式和推理参数的影响。结果:表现最佳的模型在放射科报告中提取BT-RADS评分的准确率超过98%,在病理科报告中提取IDH突变状态的准确率超过90%。表现最优的模型为医学领域微调的Llama3。更大、更新、领域微调的模型始终优于旧版和小型模型。模型量化对性能影响极小。少量示例提示显著提升准确率。RAG在复杂病理科报告中提升性能,但对较短的放射科报告无明显增益。结论:开源大语言模型在本地隐私保护环境下,展现出从非结构化临床报告中自动提取结构化数据的巨大潜力。模型选择、提示工程及基于标注数据的半自动化优化对实现最优性能至关重要。这些方法足够可靠,可用于科研工作流,凸显人机协作在医疗数据提取中的前景。

原文摘要 · Abstract (English)

Purpose: To develop and evaluate an automated system for extracting structured clinical information from unstructured radiology and pathology reports using open-weights large language models (LMs) and retrieval augmented generation (RAG), and to assess the effects of model configuration variables on extraction performance. Methods and Materials: The study utilized two datasets: 7,294 radiology reports annotated for Brain Tumor Reporting and Data System (BT-RADS) scores and 2,154 pathology reports annotated for isocitrate dehydrogenase (IDH) mutation status. An automated pipeline was developed to benchmark the performance of various LMs and RAG configurations. The impact of model size, quantization, prompting strategies, output formatting, and inference parameters was systematically evaluated. Results: The best performing models achieved over 98% accuracy in extracting BT-RADS scores from radiology reports and over 90% for IDH mutation status extraction from pathology reports. The top model being medical fine-tuned llama3. Larger, newer, and domain fine-tuned models consistently outperformed older and smaller models. Model quantization had minimal impact on performance. Few-shot prompting significantly improved accuracy. RAG improved performance for complex pathology reports but not for shorter radiology reports. Conclusions: Open LMs demonstrate significant potential for automated extraction of structured clinical data from unstructured clinical reports with local privacy-preserving application. Careful model selection, prompt engineering, and semi-automated optimization using annotated data are critical for optimal performance. These approaches could be reliable enough for practical use in research workflows, highlighting the potential for human-machine collaboration in healthcare data extraction.

医疗文本大模型信息抽取RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。