用大模型分析临床病历中的症状,提升心脏病风险预测准确率
LLM-Augmented Symptom Analysis for Cardiovascular Disease Risk Prediction: A Clinical NLP
- 用领域适配的大模型从自由文本中提取症状并推理关联
- 在MIMIC-III和CARDIO-NLP数据集上各项指标均优于传统方法
- 适合临床医生、医疗AI研发者参考,尤其关注病历智能分析
及时识别和精准评估心血管疾病(CVD)风险对降低全球死亡率至关重要。现有预测模型主要依赖结构化数据,而未结构化的临床记录中隐藏着早期预警信号。本研究提出一种基于大语言模型的临床NLP流程,利用领域适配的大模型进行症状抽取、上下文推理与关联分析。方法融合心血管领域微调、提示工程推理与实体感知推理。在MIMIC-III和CARDIO-NLP数据集上的评估显示,该方法在精确率、召回率、F1分数和AUROC上均有提升,且经心脏病专家评估具有高临床相关性(kappa = 0.82)。通过提示工程和混合规则验证,有效缓解了上下文幻觉与时间顺序模糊等挑战。本工作凸显了大模型在临床决策支持系统中的潜力,推动患者主诉向可行动的风险评估转化。
原文摘要 · Abstract (English)
Timely identification and accurate risk stratification of cardiovascular disease (CVD) remain essential for reducing global mortality. While existing prediction models primarily leverage structured data, unstructured clinical notes contain valuable early indicators. This study introduces a novel LLM-augmented clinical NLP pipeline that employs domain-adapted large language models for symptom extraction, contextual reasoning, and correlation from free-text reports. Our approach integrates cardiovascular-specific fine-tuning, prompt-based inference, and entity-aware reasoning. Evaluations on MIMIC-III and CARDIO-NLP datasets demonstrate improved performance in precision, recall, F1-score, and AUROC, with high clinical relevance (kappa = 0.82) assessed by cardiologists. Challenges such as contextual hallucination, which occurs when plausible information contracts with provided source, and temporal ambiguity, which is related with models struggling with chronological ordering of events are addressed using prompt engineering and hybrid rule-based verification. This work underscores the potential of LLMs in clinical decision support systems (CDSS), advancing early warning systems and enhancing the translation of patient narratives into actionable risk assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。