用大模型推理能力从病历中提取社会健康因素,效果媲美复杂模型。
Using reasoning LLMs to extract SDOH events from clinical notes

- 设计提示词+少量样例+自我一致性验证,提升提取准确率
- 微平均F1达0.866,性能接近顶尖模型
- 无需复杂训练,适合临床数据快速处理
健康的社会决定因素(SDOH)指影响个体生活、工作与衰老的环境、行为及社会条件,对健康结果有重大影响。然而,这些信息多存在于电子病历的非结构化临床笔记中,难以直接用于机器分析。以往研究采用基于BERT的NLP模型,虽表现良好但需复杂实现和大量算力。本研究探索具备推理能力的大语言模型(LLM)在提取结构化SDOH事件中的应用,提出四步方法:1)设计结合指南的简洁提示词;2)使用精心挑选的少量示例进行少样本学习;3)通过自一致性机制保障输出稳定性;4)后处理进行质量控制。实验结果显示,该方法微平均F1得分为0.866,性能与领先模型相当。结果表明,具备推理能力的LLM是提取SDOH事件的有效方案,兼具实现简便与优异性能。
原文摘要 · Abstract (English)
Social Determinants of Health (SDOH) refer to environmental, behavioral, and social conditions that influence how individuals live, work, and age. SDOH have a significant impact on personal health outcomes, and their systematic identification and management can yield substantial improvements in patient care. However, SDOH information is predominantly captured in unstructured clinical notes within electronic health records, which limits its direct use as machine-readable entities. To address this issue, researchers have employed Natural Language Processing (NLP) techniques using pre-trained BERT-based models, demonstrating promising performance but requiring sophisticated implementation and extensive computational resources. In this study, we investigated prompt engineering strategies for extracting structured SDOH events utilizing LLMs with advanced reasoning capabilities. Our method consisted of four modules: 1) developing concise and descriptive prompts integrated with established guidelines, 2) applying few-shot learning with carefully curated examples, 3) using a self-consistency mechanism to ensure robust outputs, and 4) post-processing for quality control. Our approach achieved a micro-F1 score of 0.866, demonstrating competitive performance compared to the leading models. The results demonstrated that LLMs with reasoning capabilities are effective solutions for SDOH event extraction, offering both implementation simplicity and strong performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。