用任务提示统一症状提取与编码,提升临床文本准确性
Task as Context Prompting for Accurate Medical Symptom Coding Using Large Language Models
- 将任务上下文嵌入提示,统一症状抽取与标准编码
- 在VAERS数据上,相比传统方法准确率显著提升
- 适合需要精准医学编码的药监与临床研究场景
从非结构化临床文本(如疫苗安全报告)中准确进行症状编码是药物警戒与安全监测的关键任务。本研究中的症状编码指将细微症状提及关联到标准化术语库(如MedDRA),区别于广义医疗编码。传统方法将抽取与链接分步处理,难以应对临床叙述的复杂性,尤其对罕见病例表现不佳。近期大语言模型虽具潜力,但性能不稳定。为此,我们提出任务作为上下文(TACO)提示框架,通过在提示中嵌入任务特定上下文,统一抽取与链接任务。研究还构建了基于VAERS报告的人工标注数据集SYMPCODER,以及两阶段评估框架,全面评估症状链接与提及保真度。对Llama2-chat、Jackalope-7b、GPT-3.5 Turbo、GPT-4 Turbo和GPT-4o等多模型的综合评估表明,TACO显著提升针对症状编码等定制任务的灵活性与准确性,为更精细编码任务及临床文本处理方法的发展铺路。
原文摘要 · Abstract (English)
Accurate medical symptom coding from unstructured clinical text, such as vaccine safety reports, is a critical task with applications in pharmacovigilance and safety monitoring. Symptom coding, as tailored in this study, involves identifying and linking nuanced symptom mentions to standardized vocabularies like MedDRA, differentiating it from broader medical coding tasks. Traditional approaches to this task, which treat symptom extraction and linking as independent workflows, often fail to handle the variability and complexity of clinical narratives, especially for rare cases. Recent advancements in Large Language Models (LLMs) offer new opportunities but face challenges in achieving consistent performance. To address these issues, we propose Task as Context (TACO) Prompting, a novel framework that unifies extraction and linking tasks by embedding task-specific context into LLM prompts. Our study also introduces SYMPCODER, a human-annotated dataset derived from Vaccine Adverse Event Reporting System (VAERS) reports, and a two-stage evaluation framework to comprehensively assess both symptom linking and mention fidelity. Our comprehensive evaluation of multiple LLMs, including Llama2-chat, Jackalope-7b, GPT-3.5 Turbo, GPT-4 Turbo, and GPT-4o, demonstrates TACO's effectiveness in improving flexibility and accuracy for tailored tasks like symptom coding, paving the way for more specific coding tasks and advancing clinical text processing methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。