对比规则方法与大模型在医疗文本提取中的精度与适应性
Balancing Natural Language Processing Accuracy and Normalisation in Extracting Medical Insights
- 用规则引擎和大模型分别处理波兰医院电子病历
- 规则方法在年龄性别提取上更准,大模型在药物识别上更强
- 翻译会损失信息,建议融合两者优势提升实用性
从非英语临床文本中提取结构化医疗信息仍是医疗领域中的开放挑战,尤其在资源匮乏的语言环境中。本研究对比了低算力规则方法与大语言模型(LLMs)在波兰阿默里卡地区疗养院电子健康记录(EHR)中的信息抽取效果。评估内容包括患者人口统计、临床发现及处方药物的提取,并考察了文本未标准化及翻译带来的信息损失影响。结果表明,规则方法在年龄、性别提取任务中准确率更高;而大模型具有更强的适应性与可扩展性,在药物名称识别上表现更优。通过对比原始波兰语文本与英文翻译文本的效果,验证了翻译会引入信息损失。研究揭示了在医疗NLP部署中精度、标准化与计算成本间的权衡,主张采用融合规则系统精确性与大模型适应性的混合方案,为真实医院环境提供更可靠、高效的临床NLP实践路径。
原文摘要 · Abstract (English)
Extracting structured medical insights from unstructured clinical text using Natural Language Processing (NLP) remains an open challenge in healthcare, particularly in non-English contexts where resources are scarce. This study presents a comparative analysis of NLP low-compute rule-based methods and Large Language Models (LLMs) for information extraction from electronic health records (EHR) obtained from the Voivodeship Rehabilitation Hospital for Children in Ameryka, Poland. We evaluate both approaches by extracting patient demographics, clinical findings, and prescribed medications while examining the effects of lack of text normalisation and translation-induced information loss. Results demonstrate that rule-based methods provide higher accuracy in information retrieval tasks, particularly for age and sex extraction. However, LLMs offer greater adaptability and scalability, excelling in drug name recognition. The effectiveness of the LLMs was compared with texts originally in Polish and those translated into English, assessing the impact of translation. These findings highlight the trade-offs between accuracy, normalisation, and computational cost when deploying NLP in healthcare settings. We argue for hybrid approaches that combine the precision of rule-based systems with the adaptability of LLMs, offering a practical path toward more reliable and resource-efficient clinical NLP in real-world hospitals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。