用本地部署的大模型自动提取病历文本中的结构化信息,提升临床研究效率。
Leveraging LLMs for Structured Data Extraction from Unstructured Patient Records
- 基于本地LLM与RAG技术,安全提取电子病历中的结构化数据。
- 在多类临床特征上表现高准确率,识别出人工漏检的错误。
- 适合需高效、一致数据采集的临床研究团队使用。
人工病历审查仍是临床研究中耗时且资源密集的环节,需专家从非结构化的电子健康记录(EHR)文本中提取复杂信息。本文提出一种安全、模块化的框架,利用本地部署的大语言模型(LLMs)在符合HIPAA标准的机构计算基础设施上实现自动化结构化特征提取。该系统将检索增强生成(RAG)与结构化输出方法集成至可广泛部署的容器中,支持多种临床领域。评估显示,该框架在大量患者病历中对多种医学特征的提取达到高准确率,相较专家标注数据集表现优异,并发现了人工审查中遗漏的若干标注错误。结果表明,该框架具备降低人工病历审查负担、提升数据采集一致性潜力,从而加速临床研究进程。
原文摘要 · Abstract (English)
Manual chart review remains an extremely time-consuming and resource-intensive component of clinical research, requiring experts to extract often complex information from unstructured electronic health record (EHR) narratives. We present a secure, modular framework for automated structured feature extraction from clinical notes leveraging locally deployed large language models (LLMs) on institutionally approved, Health Insurance Portability and Accountability Act (HIPPA)-compliant compute infrastructure. This system integrates retrieval augmented generation (RAG) and structured response methods of LLMs into a widely deployable and scalable container to provide feature extraction for diverse clinical domains. In evaluation, the framework achieved high accuracy across multiple medical characteristics present in large bodies of patient notes when compared against an expert-annotated dataset and identified several annotation errors missed in manual review. This framework demonstrates the potential of LLM systems to reduce the burden of manual chart review through automated extraction and increase consistency in data capture, accelerating clinical research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。